Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use

https://images.ctfassets.net/jdtwqhzvc2n1/TEZ90oDEVagcCeMlphUJp/c305312054f50210a953952ab93385b1/ChatGPT_Image_Aug_3__2026__07_32_45_PM.png?w=800&q=75

Chinese e-commerce and cloud giant Alibaba's famed Qwen team of AI researchers last night unveiled Qwen3.8-Max, a new flagship 2.4-trillion-parameter mixture-of-experts (MoE) multimodal large language model (LLM) that targets one of the most competitive corners of the frontier AI market: autonomous software engineering and long-horizon enterprise work.

If the company's published benchmarks hold up under broader independent testing, Qwen3.8-Max doesn't merely compete with today's leading proprietary models — it surpasses several of them on some key benchmarks in agentic computing.

Most notably, Qwen reports that Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark measuring how well ahead of GPT-5.6 Sol Max (83.2) and Fable 5 (85.0), while also posting the highest reported score on PaperBench and leading or remaining highly competitive across software engineering, research reproduction, multimodal reasoning, and visual web development benchmarks.

The release also signals a potentially significant strategic shift for Alibaba: the company says open weights for...

Copyright of this story solely belongs to venturebeat.com. To see the full text click HERE