GLM-5.3-Flash will likely handle 45% of your AI workloads

https://images.ctfassets.net/jdtwqhzvc2n1/4032PIHdBJBW2kqNHAOME0/53c12ae6fbb334fab937a4dcc1065c33/Gemini_Generated_Image_3lw7jd3lw7jd3lw7.jpeg?w=800&q=75

A week ago, a mystery model called Ox Alpha showed up on OpenRouter — one more entrant among more than 400 models, with roughly 10 new ones launching every week. What made it stand out wasn't just the free price tag; it was quietly good. Hobbyists and indie developers noticed fast, pushing several trillion tokens through it daily, with community estimates for the week ranging from single digits to over 20 trillion.

AI enthusiasts spent the next six days doing forensics and speculating who could have built it, and who could have the infrastructure to serve that many tokens for free. First the guess was a U.S. lab: the long-awaited Gemini, or Anthropic shipping a good-enough middle tier, or Elon sitting on so much capacity he dropped Ox Alpha (note the naming). People ran tokenizer traces and networking analysis. A real Sherlock Holmes mystery week.

On August 26, Z.ai put...

Copyright of this story solely belongs to venturebeat.com. To see the full text click HERE

Read more