Jonathan Ross

@jonathan-ross.bsky.social

CEO + Founder @ Groq, the Most Popular API for Fast Inference | Creator of the TPU and LPU, Two of the World’s Most Important AI Chips | On a Mission to Double the World's AI Compute by 2027

Founder Tip #2: You have to spend time to make time. Hiring, re-organizing, calendar clean up (across the team), preparation for meetings (internal and external), etc. Half my day is available for whatever I find important - because the other half is spent freeing up time.

Clearly China doesn't have enough compute for scaled AI today: - GPT-OSS, Llama [US]: optimized for cheaper inference - R1, Kimi K2, Qwen [China]: optimized for cheaper training With China's population reducing inference costs is more important, and that means more training.

I spent the weekend hanging out with a group of friends. A question we asked was what dreams did we have that we gave up on? When I was 18, I had two dreams: 1) Be an astronaut 2) Build AI chips I didn’t give up on one of them. 😀

Big news! Mistral AI Saba 24B is on GroqCloud! The specialized regional language model is perfect for Middle East and South Asia-based devs and enterprises building AI solutions that need fast inference. Learn more: groq.com/mistral-saba...

Mistral Saba Added to GroqCloud™ Model Suite - Groq is Fast AI Inference

GroqCloud™ has added another openly-available model to our suite – Mistral Saba. Mistral Saba is Mistral AI’s first specialized regional language model,

hubs.la

It was a pleasure being back on 20VC with Harry Stebbings. His craft of interviewing is second to none and we went deep. This is the interview after we just launched 19,000 LPUs in Saudi Arabia. We built the largest inference cluster in the region. Link to the interview in the comments below!

When you make compute cheaper do people buy more? Yes. It's called Jevons Paradox and it's a big part of our business thesis. In the 1860s, an Englishman wrote a treatise on coal where he noted that every time steam engines got more efficient people bought more coal. 🧵(1/5)

Bild

This is insane, Groq is the #4 API on this list! 😮 OpenAI, Anthropic, and Azure are the top 3 LLM API providers on LangChain Groq is #4, and close behind Azure Google, Amazon, Mistral, and Hugging Face are the next 4. Ollama is for local development. Now add three more 747's worth of LPUs 😁

Bild

(1/5) One of the reasons why chips are so hard to innovate in is because if you're asking someone to put up a 10 million, 100 million, or a billion dollar check they need to know that what they're buying is going to work.

Bild

(1/5) Everyone at Groq has one of these challenge coins on them. It’s how we create alignment. One side says its 25 million, because we're going to get to 25 million tokens per second by the end of the year On the other side, it says, “Make it real. Make it now. Make it wow.”

Bild

What can you do with Llama quality and Groq speed? Instant. That's what. 3 months back: Llama 8B running at 750 Tokens/sec Now: Llama 70B model running at 3,200 Tokens/sec We're still going to get a liiiiiiitle bit faster, but this is our V1 14nm LPU - how fast will V2 be? 😉