Qwen 4 · Research

Alibaba Confirms Qwen 4 in Training as Apsara Roadmap Targets 10T Parameters

Hub-and-spoke data poster titled Qwen 4 Confirmed In Training, with satellites for Qwen 4.5, Qwen 5, the Zhenwu V900 accelerator and Alibaba Cloud capacity targets.
AK

Threat intelligence editor · Updated Sep 22, 2026, 5:28 AM EDT

Alibaba confirmed Qwen 4 is in training at Apsara 2026, with Qwen 4.5 and Qwen 5 targeting 5-10T parameters and the Zhenwu V900 chip due Q1 2027.

Alibaba used the opening day of its 2026 Apsara Conference in Hangzhou to confirm what the open-weight community has been reading tea leaves about since July: Qwen 4 exists, and it is currently in training.

That is the whole of the confirmation. There is no Qwen 4 release, no API endpoint, no published parameter count, and no weights on Hugging Face. What Alibaba delivered instead was a roadmap — and the roadmap is the story, because it puts a number on where the Qwen line is going and pairs it with the silicon and the power to get there.

What Alibaba actually confirmed

Three commitments came out of the Apsara keynote, and they are worth separating from the speculation that has attached itself to them.

Qwen 4 is in training. Alibaba's own wording is that its "next-generation model, Qwen 4, is currently in training." No release window was given. No size was disclosed.

Qwen 4.5 and Qwen 5 are on the roadmap, targeting 5 to 10 trillion parameters. This is the first time Alibaba has publicly committed to a parameter ceiling beyond the current generation. For calibration, the largest model Alibaba ships today is Qwen3.8-2.4T-A95B — 2.4 trillion total parameters with 95 billion active. A 10T target is roughly a 4x jump in total scale from a model that only shipped in August.

The Zhenwu V900 gives it somewhere to run. Alibaba unveiled its next-generation accelerator alongside the model roadmap: 216 GB of on-package memory, 1,200 GB/s of inter-chip bandwidth, and a claimed 3x the performance of the Zhenwu M890 it released in May. Mass production and commercial availability are scheduled for Q1 2027 — which lines up suggestively with a Qwen 4 that is in training now.

Underneath all of it sits the capacity commitment: Alibaba Cloud intends to expand global datacenter capacity past 20 gigawatts by 2032. Markets liked the package; Alibaba shares rose roughly 3% in Hong Kong on the announcement.

The architecture is already public

The most useful thing about the Qwen 4 announcement is that you do not have to wait for Qwen 4 to see how it will be built. Alibaba shipped the architecture preview a month ago.

Qwen3.8-Flash-Next, released on 24 August, was explicitly positioned by the Qwen team as a preview of the next-generation architecture rather than a flagship in its own right. Its shape is the tell:

  • A mixture-of-experts backbone of roughly 125 billion parameters, supplemented by an additional 51 billion n-gram embedding parameters
  • Only about 6 billion parameters activated per token
  • Native multimodal support

That activation ratio is the signal. Qwen3.8-Max, the current hosted flagship, runs 2.4T total against 95B active — a sparsity ratio around 25:1. Flash-Next runs roughly 176B total against 6B active, closer to 29:1, while pushing a meaningful share of its capacity into n-gram embeddings that cost lookup rather than compute.

Extrapolate that to a 10T-parameter Qwen 5 and the active-parameter count does not have to grow anything like as fast as the total. That is the entire argument for why a 10T target is a serving proposition rather than a vanity metric, and it is why the Flash-Next preview matters more than the headline number.

Architecture diagram comparing total and active parameter counts across Qwen3.8-27B, Qwen3.8-Flash-Next and Qwen3.8-Max, against the 5-10T target for Qwen 5.

Total parameters keep climbing; active parameters per token do not. Flash-Next is the architecture Qwen 4 is expected to inherit.

What was not confirmed

A widely-circulated social media post claimed that a full Qwen 4 lineup — Qwen-4-Max, Qwen-4-Flash, Qwen-4-Plus, and a dense Qwen-4-27B — was officially announced at Apsara. Treat that as unconfirmed.

It traces to a single account. It does not appear in the mainstream conference coverage, which reports only that Qwen 4 is in training. And it is contradicted by the most checkable source available: as of 22 September 2026, the Qwen organisation on Hugging Face contains no model named Qwen4 or Qwen-4. The most recent entries are Qwen-Image-2.1-PE (20 September), Qwen-Image-2.1 (14 September), Qwen-Drive-1.0-4B (27 August), and Qwen3.8-Flash-Next (24 August).

The lineup is plausible — it mirrors the Qwen3.8 structure of a hosted Max tier, Flash and Plus service tiers, and an open-weight dense model around 27B — but plausible is not announced. If a Qwen-4-27B lands, it will land on Hugging Face, and that is where to verify it.

One further claim from the conference deserves the same caution. Alibaba presented results it attributes to recursive self-improvement: over roughly a month of fully automated runs, Qwen3.8-Max is said to have completed 33 iterative cycles, raising its Artificial Analysis score from 40 to 45 through autonomous training optimisation. That is a vendor-reported figure from a keynote, not an independently reproduced result.

Why the 27B tier is the one to watch

For anyone running models on their own hardware, the interesting question is not whether Qwen-4-Max beats the frontier. It is whether the open-weight dense tier survives the generation.

Alibaba has been unusually consistent here. Qwen3.6-27B shipped in April. Qwen3.8-27B shipped in August, with an FP8 variant a week later. A 27B dense model is the largest thing that fits comfortably in 24 GB of VRAM at 4-bit quantisation and runs at usable speed on a single consumer card or a mid-tier Apple Silicon machine — the size class that carries most serious local deployment.

Nothing Alibaba said at Apsara commits it to continuing that tier into Qwen 4. The roadmap it did announce points the other way, toward trillion-parameter hosted models running on proprietary accelerators that nobody outside Alibaba Cloud will operate. The 10T headline and the 27B open-weight release are not in tension today, but they are pulling in different directions, and the Qwen 4 launch is where that tension gets resolved.

What this means in practice

Nothing changes in your stack this week. Qwen 4 is in training, not in production, and the accelerator that is presumably meant to serve it does not reach mass production until Q1 2027.

What has changed is the planning horizon:

  • If you are standardising on Qwen3.8-27B, you have a stable target. The architecture preview is already public, and a Qwen 4 dense model would be an upgrade path, not a migration.
  • If you are building against the hosted Max tier, expect the sparsity ratio to keep widening. Cost per token has been falling on active-parameter efficiency, not on discounting, and Flash-Next suggests that trend continues.
  • If your procurement depends on open weights, the Qwen 4 launch is the checkpoint. Alibaba's open-weight commitment has been real but has never been contractual, and a 10T roadmap gives it a reason to reconsider.

The right move is to watch Hugging Face rather than the keynotes. Alibaba has shipped its real architecture disclosures as model cards, a month ahead of the announcements — and it will almost certainly do it again.