What Nvidia's first Groq 3 LPU benchmarks do and don't tell us about its $20B gamble
Nvidia’s $20 billion bet on Groq’s LPU tech sure looks like it was a good one. On Monday, the GPU giant offered the first glimpse of just how big a speedup its Groq 3-based LPX racks will provide.
In an independent benchmark conducted by Artificial Analysis, Nvidia’s LPX rack systems managed to churn out 3,400 tokens a second (tok/s) with a 100,000-token input sequence in Google’s Gemma 4 31B model.
According to Nvidia, this makes it 4x faster than the nearest alternative platform, which going off Artificial Analysis’ leaderboard would be a direct dig at Cerebras, which managed a still impressive 882 tok/s under the same conditions.
Acquihiredby Nvidia in late December, Groq has LPUs that feature an SRAM-heavy dataflow architecture designed specifically for high-performance inference serving. Unlike traditional datacenter GPUs, which rely on high-speed DRAM memory tech like GDDR7 and HBM4, Groq’s chips rely entirely on a large...
Copyright of this story solely belongs to theregister.com. To see the full text click HERE