Qwen 3.8-Max, Opus 5 show why scores miss cost | VentureBeat

https://images.ctfassets.net/jdtwqhzvc2n1/6zpX2RqBlYloDvRCIfyTbk/53362a14bf8566e0ca2a910e5af9f6e4/Gemini_Generated_Image_x0e43ax0e43ax0e4.png?w=800&q=75

Alibaba released Qwen 3.8-Max this week and marketed the preview as second only to Claude Fable 5 (their launch-day table was more equivocal: the model leads on one of 12 coding-agent rows). But an independent harness came close to the opposite conclusion: a benchmark run, apparently using the Preview version, put Qwen 3.8-Max's best effort setting mid-pack, and its default setting last.

Both results are real and defensible. The gap between them is about token and time budgets, and that matters because those figures aren’t usually headline numbers. Alibaba's footnotes give its coding numbers a five-hour timeout, and up to 12 hours per run on PaperBench. The independent harness, VulcanBench, allowed between 45 and 60 minutes of wall clock time. A time budget between five and 16 times larger on Alibaba’s side explains the huge difference in results.

It’s time to do two things to start accounting for...

Copyright of this story solely belongs to venturebeat.com. To see the full text click HERE

Read more