Alibaba's 2.4 trillion parameter Qwen3.8-Max arrived this week as Chinese systems claimed up to 46% of U.S. enterprise token traffic, squeezing OpenAI's GPT-5.6.
Key Points:
- Alibaba's Qwen3.8-Max carries 2.4 trillion total parameters and activates 95 billion per request, with open weights promised next week.
- GPT-5.6 Sol still leads the agentic coding benchmark Alibaba published, but the margins are narrow and unverified by outside labs.
- OpenAI cut its cheapest tier by 80% on Jul. 30, three weeks after launch, as Chinese models took volume.
Qwen3.8-Max Benchmarks Test GPT-5.6 Sol
Alibaba unveiled the model on Aug. 3, a mixture-of-experts system that carries 2.4 trillion total parameters but activates only 95 billion of them on any single request.
The company said open weights for the flagship and a smaller 27-billion-parameter version would reach public repositories next week, reversing a run of closed releases from its top model line. Alibaba also published a full benchmark table.
Those figures put Qwen3.8-Max at 86.6 on Terminal-Bench 2.1, an agentic coding test where OpenAI's GPT-5.6 Sol still leads the field at 88.8. On SWE-bench Pro it reached 67.7 against 80.0 for Anthropic's Claude Fable 5, and on OSWorld-Verified, a computer-use test, it claimed 86.1 versus Sol's 83.2.
No independent lab has reproduced the results. Alibaba said the model spent 16 days building a command-line tool without human help, turning incoming user requests into working code and running its own tests. Developers questioned the demonstration and the license terms.
Also Read: iPhone 18 Pro Max Could Be First With Apple's 2nm Chip, 15% Faster
Ion Stoica Weighs China's Narrowing Gap
Ion Stoica, a University of California, Berkeley computer scientist who co-founded the Arena benchmark platform, assessed in late July that Chinese open-weight models now trail American frontier labs by two or three months. He had earlier put that distance at six to nine months, a shift that tracks a year of open releases from DeepSeek and Z.ai's GLM line.
"Price is doing the work here," said Harpreet Arora, who runs agentic infrastructure at Vercel, describing engineering teams that route ordinary work to whichever cheap model clears the bar. Justin Summerville of OpenRouter measured open Chinese systems at 60% to 90% cheaper. The routing follows the math.
GPT-5.6 Pricing Defends The Volume Tier
OpenAI cut its Luna tier by 80% on Jul. 30, three weeks after the GPT-5.6 family launched, to 20 cents per million input tokens and $1.20 per million output tokens. Terra fell 20%, while flagship Sol held at $5 and $30.
That pattern keeps a premium price at the top of the family while the cheaper tiers fight for volume Chinese labs have been taking all year.
Chinese-origin models climbed from 4.5% of U.S. token volume on one large routing platform in the first half of 2025 to more than 30% every week since February, peaking at 46%. Stanford University researchers recorded the top American model's benchmark lead at 2.7% in March, against as much as 31.6 percentage points in 2023.
Read Next: Bitcoin Could Turn Extreme Fear Into Fuel For A $75K Rally






