Nvidia and Cerebras are selling performance their customers will (probably) never see

https://image.theregister.com/251642.jpg?imageId=251642&x=0&y=0&cropw=100&croph=100&panox=0&panoy=0&panow=100&panoh=100&width=1200&height=683

Shots were fired at the Hot Chips conference in California this week as Nvidia announced its new Groq-3-based LPX racks had entered production, with early tests showing the systems churning an eye-watering 3,400 tokens a second in Gemma 4 31B. That's four times faster than rival Cerebras.

A day later, Cerebras fired back, touting nearly equivalent performance from its next-gen CS-4 accelerators revealed last week.

These top-line performance figures make inference feel instantaneous relative to the chatbots we've grown accustomed to over the past four years. But the two companies are essentially arguing over numbers their customers will probably never see in production.

To be clear, neither is lying. If you wanted to recreate these results, you certainly could — Artificial Analysis, the team responsible for both sets of benchmarks, knows what they're doing — but beyond a marketing gimmick, no inference-as-a-service model operator in their right mind would run...

Copyright of this story solely belongs to theregister.com. To see the full text click HERE

Read more