What Nvidia's first Groq 3 LPU benchmarks do and don't tell us about its $20B gamble

https://image.theregister.com/5224053.jpg?imageId=5224053&x=0&y=0&cropw=100&croph=100&panox=0&panoy=0&panow=100&panoh=100&width=1200&height=683

Nvidia’s $20 billion bet on Groq’s LPU tech sure looks like it was a good one. On Monday, the GPU giant offered the first glimpse of just how big a speedup its Groq 3-based LPX racks will provide.

In an independent benchmark conducted by Artificial Analysis, Nvidia’s LPX rack systems managed to churn out 3,400 tokens a second (tok/s) with a 100,000-token input sequence in Google’s Gemma 4 31B model.

According to Nvidia, this makes it 4x faster than the nearest alternative platform, which going off Artificial Analysis’ leaderboard would be a direct dig at Cerebras, which managed a still impressive 882 tok/s under the same conditions.

Acquihiredby Nvidia in late December, Groq has LPUs that feature an SRAM-heavy dataflow architecture designed specifically for high-performance inference serving. Unlike traditional datacenter GPUs, which rely on high-speed DRAM memory tech like GDDR7 and HBM4, Groq’s chips rely entirely on a large...

Copyright of this story solely belongs to theregister.com. To see the full text click HERE

Read more