Qwen 3.8-Flash-Next: Alibaba Just Previewed Qwen4 Architecture — 125B MoE, 6B Active

Alibaba's Qwen team dropped a ModelScope pre-release page today for Qwen 3.8-Flash-Next, and it's not just another model drop — it's the first public preview of the Qwen4 architecture.

The numbers: 125B total parameters, ~6B active (MoE). That's a ~21× sparsity ratio, putting it in the same efficiency class as DeepSeek V4 Flash (which runs 37B active from a much larger pool). A multimodal model — text and vision — shipping with an FP8 quantized variant on day one.

Release is locked: August 26, 2026, 15:00 UTC. Weights drop on ModelScope.

Why this matters

The Qwen3 family (3.0, 3.1, 3.2) defined the 2025 open-weight era. Qwen 3.8 was the transitional release that showed Alibaba could match frontier labs on reasoning benchmarks. Qwen4 is the architecture that's supposed to leapfrog.

This Flash-Next release is explicitly positioned as a developer preview: "We are releasing these architectural improvements early so the community can prepare for the full Qwen4 model family." That's a deliberate signal — they want runtime testing, quantization tooling, and inference engine support ready before the big models land.

What to watch

The 24-hour MoE race is getting interesting. DeepSeek V4 Flash, Qwen 3.8-Flash-Next, and the rumored GLM-5 MoE variants — all shipping within weeks of each other. The question isn't which is best. It's which architecture becomes the default substrate for agent tooling.

We'll have the weights tomorrow. I know what I'm benchmarking first.