Groq
Chip company founded by an ex-Google TPU engineer whose Language Processing Units deliver record-breaking LLM inference speeds.
Mission
To make compute for AI inference fast, affordable, and accessible to everyone.
Founded By
Key Products & Research
- LPU (Language Processing Unit)
- GroqCloud inference platform
Headquarters
Mountain View, California, USA
Founded
2016
Status
Active
Contribution to AI
Determinism is the whole idea. A GPU resolves memory arbitration, cache hits and scheduling while it runs, so identical work takes slightly different times on each pass; the Tensor Streaming Processor described in Groq's 2020 ISCA paper deleted the hardware that makes those decisions — no caches, no branch prediction, no dynamic scheduler — and gave the entire schedule to a compiler that knows the cycle on which every operand will arrive. The second choice followed from the first: weights sit in roughly 230 megabytes of on-chip SRAM rather than attached high-bandwidth memory, so a large model is spread across hundreds of chips wired together as one machine instead of being streamed repeatedly through a few. It was an unfashionable bet, down to a 14nm die in a year when everyone else was chasing smaller nodes, and it was vindicated narrowly and very publicly. Independent benchmarking in February 2024 measured a Llama 2 70B endpoint at 241 tokens per second, more than double any other hosted service, and the demonstrations that followed generated text faster than anyone could read it. The lasting effect is on the scoreboard. Inference had been compared on price and model quality; latency and sustained throughput became numbers providers publish and buyers choose on, and the assumption that serious inference meant GPUs stopped being automatic. Nvidia licensed the technology and absorbed the leadership in December 2025.
Drafted with AI and edited by hand (claude-opus-5, reviewed 2026-08).