Qwen3.8-Max Draws Level With the Frontier on Agentic Benchmarks

Alibaba released Qwen3.8-Max on August 3, 2026, and within days independent evaluation put it level with the frontier on agentic work. On Artificial Analysis’ Agentic Index, the 2.4-trillion-parameter model scores 58 — tied with Claude Opus 5 running at xhigh effort, and one point behind Opus 5 at max effort, which still holds the top spot at 59. It is the closest a Chinese lab has come to the top of that particular leaderboard, and the weights are due to be published.
Intermediate
What the Model Is
Qwen3.8-Max is a sparse mixture-of-experts model with 2.4 trillion total parameters that activates roughly 95 billion per token — about a 4% activation ratio. It takes text, image, and video input, returns text, and carries a 1-million-token context window with up to 128k tokens of output. API pricing is $2.00 per million input tokens and $6.00 per million output tokens, with cache hits at $0.25 — a cut from the $2.50/$7.50 the preview generation charged.
The Qwen team’s own framing at announcement was that it is “one of the most powerful model available today, compatible to leading frontier AI models, second only to Fable 5.” Alibaba’s self-reported benchmarks put it at 86.1 on OSWorld-Verified — ahead of GPT-5.6 Sol Max at 83.2, Claude Fable 5 at 85.0, and Gemini 3.1 Pro at 76.2 — along with 86.6 on Terminal-Bench 2.1, 67.7 on SWE-bench Pro, 92.6 on GPQA Diamond, and 93.0 on PaperBench. Those are vendor-run numbers using each rival’s own coding harness, and worth treating as such until independently reproduced.
What Independent Testing Found
On the broader Artificial Analysis Intelligence Index — version 4.1, which weights GDPval-AA v2 at 20%, Terminal-Bench 2.1 at 16%, and τ³-Bench Banking at 14%, alongside Humanity’s Last Exam, SciCode, GPQA, CritPt, AA-LCR, and AA-Omniscience — Qwen3.8-Max landed at 56 at launch, level with Claude Opus 4.8 and one point behind Kimi K3 at 57. Artificial Analysis has since revised the figure upward to 58. The number moved twice: an initial score of 53 was withdrawn after what Artificial Analysis described as intermittent issues on the endpoint being tested.
The agentic gains are real but expensive. On GDPval-AA, which runs models through tasks drawn from 44 occupations with shell access and web browsing, Qwen3.8-Max posts 1,739 Elo — a 468-point jump over its predecessor, ahead of GPT-5.6 Sol Max at 1,730 and Kimi K3 at 1,685, behind Claude Opus 5 at 1,852. It gets there by working much harder: Artificial Analysis measured it taking 64 turns per task, against 14 for Qwen3.7-Max. Cost per Intelligence Index task came to $1.14, more than double the previous generation’s $0.53 and above Kimi K3’s $0.86.
Two scores went backwards. AA-LCR dropped two points, and AA-Omniscience fell ten, with the measured hallucination rate rising from 23% to 40%.
What This Means
The headline result is not that Qwen3.8-Max beat Opus 5 — it did not — but that the gap at the top of the agentic leaderboard is now roughly one index point, and the model sitting there is one whose weights Alibaba says it will publish. This would be the first Max-class Qwen released openly, distributed through Alibaba Cloud’s Model Studio alongside a smaller Qwen3.8-27B.
How much that openness is worth in practice is a separate question. A 2.4T checkpoint is a multi-node datacentre artifact; running it is out of reach for individual developers and most companies. Amit Jena of Kanerika made the narrower point that the release is still an intention rather than a fact: “Publishing weights is a separate act from opening an API endpoint. Until there is a repository, a licence and a model card, open-weight describes an intention.” Nitish Tyagi, a senior principal analyst at Gartner, noted that organisations outside China may hesitate to depend on models hosted within it, and that open-weight models typically lack the indemnification that commercial vendors provide.
Alibaba also says the model completed a software engineering project autonomously over 16 days — a claim that, as Jena observed, arrives without the detail that would make it checkable: “Sixteen days of what? How many times did a human step in? Did the output survive code review?” Charlie Dai, VP and principal analyst at Forrester, put the release in context: “Alibaba is narrowing the gap, but the larger story is the rapid maturation of open-weight models. Enterprises increasingly have credible alternatives to proprietary frontier models.”
The efficiency picture is the one worth watching. A model that reaches near-frontier agentic scores by taking four and a half times as many turns is buying capability with inference budget, and the cost-per-task chart shows exactly where that lands. Whether the open weights change that calculus depends on details — licence terms, hardware requirements, and a model card — that had not been published at the time of writing.
Related Coverage
- Qwen3.8-Max Preview: Alibaba’s 2.4T-Parameter Bid for the Frontier — our coverage of the July 19 preview announcement, two weeks before this release
- Qwen3.6-35B-A3B: Alibaba Open-Sources a Frontier-Class Agentic Coder — the April open-weight release that set the pattern this one extends
- Qwen-Image-3.0: 4.5k-Token Prompts, No Benchmarks, No Weights — a Qwen release that went the other way on openness
- Anthropic Releases Claude Opus 4.8 for Longer Agentic Coding — the model Qwen3.8-Max draws level with on the Intelligence Index
This post was drafted with AI assistance and reviewed by RITS staff.
Sources
- Artificial Analysis — Agentic Index leaderboard
- Artificial Analysis — Qwen3.8 Max model page
- Artificial Analysis — Intelligence Index v4.1 methodology
- The Decoder — Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for 25 percent less
- OfficeChai — Qwen 3.8 Max scores 56 on Artificial Analysis Intelligence Index
- InfoWorld — Alibaba takes aim at OpenAI and Anthropic with Qwen3.8-Max launch
- Latent Space AINews — Qwen 3.8 Max (2.4T) and 27B, new open weights models




沪公网安备31011502017015号