Alibaba at Apsara 2026: Qwen 4 in Training, 10T Models Next

At its annual Apsara Conference in Hangzhou on September 22, 2026, Alibaba laid out the next stage of its model roadmap: Qwen 4 is now in training, and the Qwen 4.5 and Qwen 5 series that follow are projected to reach 5 to 10 trillion parameters. In the same keynote, Alibaba Group CEO Eddie Wu introduced the Zhenwu V900, T-Head’s next-generation AI accelerator. Wu also set a target of more than 20 GW of Alibaba Cloud data-center capacity by 2032, and Alibaba reported early results from getting Qwen3.8-Max to improve its own training pipeline.

Intermediate

Eddie Wu on stage at the 2026 Apsara Conference in front of a slide showing Alibaba Cloud's 20 GW data-center capacity target for 2032, a tenfold increase over 2022
Eddie Wu presenting Alibaba Cloud’s 20 GW capacity target for 2032. Image credit: Alibaba Cloud, via astig.ph

The Model Roadmap: Qwen 4 Now, 5–10T Next

Alibaba’s conference release says only that Qwen 4 is “currently in training”. It gives no release date, sizes, or licensing terms. The larger news is the generation after it. In his keynote, Wu said the Qwen team “plans to train a new model at the scale of 5 to 10 trillion-parameter, with the goal of completing more complex, longer-horizon tasks and advancing toward ASI.” The release attaches that scale to the Qwen 4.5 and Qwen 5 series.

For comparison, Alibaba’s current flagship, Qwen3.8-Max, has 2.4 trillion parameters, according to Reuters. The planned models would therefore be roughly two to four times larger. A post on X that was widely shared after the keynote listed Qwen-4-Max, Qwen-4-Plus, Qwen-4-Flash and Qwen-4-27B as the Qwen 4 lineup. Alibaba’s published materials do not name these variants, so treat them as unconfirmed.

Wu did single out the current open-weight 27B model, saying “our open-source Qwen-27B has emerged as the most popular model among developers worldwide.” He said Alibaba will keep refining models built for local deployment.

Recursive Self-Improvement, With Numbers

Wu called recursive self-improvement (RSI) the “concrete path forward.” In RSI, a model designs its own experiments, synthesizes data, and evaluates the outcomes. Alibaba attached figures to the idea:

  • Model training: Over a month of fully automated runs covering pipeline design, data validation, experimentation and error diagnosis, Qwen3.8-Max completed 33 iterative cycles. Its Artificial Analysis score rose from 40 to 45.
  • Chip design: In more than 60 hours of self-improvement and over 10,000 EDA tool calls, the system produced “production-grade chip bus modules” that were 42% smaller in chip area.

Alibaba has not published a technical report on these runs. For now, they are company-reported results without a disclosed methodology.

Zhenwu V900 and the Compute Build-Out

The Zhenwu V900, from Alibaba’s T-Head chip unit, succeeds the Zhenwu M890. Alibaba describes it as “the most powerful AI chip in China today.” The company’s stated specifications:

  • Performance: three times that of the M890
  • Memory: 216 GB (up from 144 GB of HBM on the M890)
  • Inter-chip bandwidth: 1,200 GB/s (up from 800 GB/s)
  • Scale: a single V900 cluster can support up to 500,000 cards for training and inference
  • Availability: mass production and commercial release expected in Q1 2027

Alibaba did not publish absolute compute figures (FLOPS at a given precision), so the “three times” claim can only be read against the M890. According to the release, the Zhenwu line currently serves more than 650 customers. Wu also said Alibaba’s M890 AI Supernode already runs inference for foundation models above 2 trillion parameters, and that supernodes go into commercial-scale deployment this quarter.

The model and chip plans depend on a larger infrastructure target. Wu said Alibaba Cloud’s global data-center capacity will surpass 20 GW by 2032. The keynote slide shows that as a tenfold increase over 2022. He also acknowledged that “global shortages across the AI data center supply chain are currently limiting the speed at which we can scale our compute infrastructure.”

Also Announced

The conference also brought a set of multimodal and device-side releases. They include Qwen3.8-LiveTranslate for simultaneous interpretation, and the Qwen-Audio-3.1 family (ASR, TTS, TTS-Next and Realtime). Alibaba also launched Qwen Intelligence, an end-to-end agent platform for smartphone makers, and said Qwen-Image 3.1 will launch later in 2026.

What This Means

The headline is a plan, not a product. Qwen 4 has no release date, and the 5–10T models are at least one generation further out. Still, the roadmap is unusually specific. Alibaba has given a parameter range for models two releases ahead, a quarter for its next accelerator, and a capacity figure for 2032. It has matched the model scale to hardware sized for it: a 500,000-card cluster is the kind of footprint a 10T-parameter training run implies.

Two open questions matter most to researchers. The first is whether Qwen 4 and its successors keep the open-weight releases that have made Qwen a default base model in open research. Alibaba said nothing about licensing at Apsara, though Wu’s praise for Qwen-27B suggests the small-model line will continue. The second is whether the RSI figures hold up under outside evaluation. A five-point Artificial Analysis gain from automated post-training would be notable if it were reproduced and documented.

Related Coverage

This post was drafted with AI assistance and reviewed by RITS staff.

Sources