Qwen3.8-Max Draws Level With the Frontier on Agentic Benchmarks

Alibaba released Qwen3.8-Max on August 3, 2026, and within days independent evaluation put it level with the frontier on agentic work. On Artificial Analysis’ Agentic Index, the 2.4-trillion-parameter model scores 58 — tied with Claude Opus 5 running at xhigh effort, and one point behind Opus 5 at max effort, which still holds the top spot at 59. It is the closest a Chinese lab has come to the top of that particular leaderboard, and the weights are due to be published.

Intermediate

Three bar charts from Artificial Analysis comparing leading models on Intelligence Index score, output speed in tokens per second, and cost per Intelligence Index task, with Qwen3.8 Max highlighted in orange
Image credit: Artificial Analysis, via The Decoder

What the Model Is

Qwen3.8-Max is a sparse mixture-of-experts model with 2.4 trillion total parameters that activates roughly 95 billion per token — about a 4% activation ratio. It takes text, image, and video input, returns text, and carries a 1-million-token context window with up to 128k tokens of output. API pricing is $2.00 per million input tokens and $6.00 per million output tokens, with cache hits at $0.25 — a cut from the $2.50/$7.50 the preview generation charged.

The Qwen team’s own framing at announcement was that it is “one of the most powerful model available today, compatible to leading frontier AI models, second only to Fable 5.” Alibaba’s self-reported benchmarks put it at 86.1 on OSWorld-Verified — ahead of GPT-5.6 Sol Max at 83.2, Claude Fable 5 at 85.0, and Gemini 3.1 Pro at 76.2 — along with 86.6 on Terminal-Bench 2.1, 67.7 on SWE-bench Pro, 92.6 on GPQA Diamond, and 93.0 on PaperBench. Those are vendor-run numbers using each rival’s own coding harness, and worth treating as such until independently reproduced.

What Independent Testing Found

Artificial Analysis Intelligence Index bar chart showing Claude Opus 5 at 61, Claude Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57 and Qwen3.8 Max at 56, above a scatter plot of Intelligence Index against cost per task on a logarithmic scale
Image credit: Artificial Analysis, via OfficeChai

On the broader Artificial Analysis Intelligence Index — version 4.1, which weights GDPval-AA v2 at 20%, Terminal-Bench 2.1 at 16%, and τ³-Bench Banking at 14%, alongside Humanity’s Last Exam, SciCode, GPQA, CritPt, AA-LCR, and AA-Omniscience — Qwen3.8-Max landed at 56 at launch, level with Claude Opus 4.8 and one point behind Kimi K3 at 57. Artificial Analysis has since revised the figure upward to 58. The number moved twice: an initial score of 53 was withdrawn after what Artificial Analysis described as intermittent issues on the endpoint being tested.

The agentic gains are real but expensive. On GDPval-AA, which runs models through tasks drawn from 44 occupations with shell access and web browsing, Qwen3.8-Max posts 1,739 Elo — a 468-point jump over its predecessor, ahead of GPT-5.6 Sol Max at 1,730 and Kimi K3 at 1,685, behind Claude Opus 5 at 1,852. It gets there by working much harder: Artificial Analysis measured it taking 64 turns per task, against 14 for Qwen3.7-Max. Cost per Intelligence Index task came to $1.14, more than double the previous generation’s $0.53 and above Kimi K3’s $0.86.

Two scores went backwards. AA-LCR dropped two points, and AA-Omniscience fell ten, with the measured hallucination rate rising from 23% to 40%.

What This Means

The headline result is not that Qwen3.8-Max beat Opus 5 — it did not — but that the gap at the top of the agentic leaderboard is now roughly one index point, and the model sitting there is one whose weights Alibaba says it will publish. This would be the first Max-class Qwen released openly, distributed through Alibaba Cloud’s Model Studio alongside a smaller Qwen3.8-27B.

How much that openness is worth in practice is a separate question. A 2.4T checkpoint is a multi-node datacentre artifact; running it is out of reach for individual developers and most companies. Amit Jena of Kanerika made the narrower point that the release is still an intention rather than a fact: “Publishing weights is a separate act from opening an API endpoint. Until there is a repository, a licence and a model card, open-weight describes an intention.” Nitish Tyagi, a senior principal analyst at Gartner, noted that organisations outside China may hesitate to depend on models hosted within it, and that open-weight models typically lack the indemnification that commercial vendors provide.

Alibaba also says the model completed a software engineering project autonomously over 16 days — a claim that, as Jena observed, arrives without the detail that would make it checkable: “Sixteen days of what? How many times did a human step in? Did the output survive code review?” Charlie Dai, VP and principal analyst at Forrester, put the release in context: “Alibaba is narrowing the gap, but the larger story is the rapid maturation of open-weight models. Enterprises increasingly have credible alternatives to proprietary frontier models.”

The efficiency picture is the one worth watching. A model that reaches near-frontier agentic scores by taking four and a half times as many turns is buying capability with inference budget, and the cost-per-task chart shows exactly where that lands. Whether the open weights change that calculus depends on details — licence terms, hardware requirements, and a model card — that had not been published at the time of writing.

Related Coverage

This post was drafted with AI assistance and reviewed by RITS staff.

Sources