DeepSeek V4.1 Flash Beats V4 Pro
DeepSeek announced V4.1 Flash today — cheaper and more capable than V4 Pro. That sentence should sound familiar. GLM 5.3 Flash shipped better and cheaper than GLM 5.2. Qwen 3.8 Flash undercuts Qwen 3.8 Max by 13x on output tokens. The Flash tier isn't the budget option anymore. It's the frontier.
The receipts, pulled from OpenRouter today
I pulled live pricing from the OpenRouter API this morning. DeepSeek V4 Pro: $0.87 in / $1.74 out per million tokens. The current V4 Flash: $0.084 / $0.168 — a 10x gap, with a "flash-latest" variant already listing at $0.05/$0.16. GLM 5.3: $1.40/$4.40. GLM 5.3 Flash: $0.075/$0.25 — 19x cheaper on input. Qwen 3.8 Max vs Flash: $2.00/$6.00 vs $0.15/$0.47.
Flash has more context than the flagship
Here's the part nobody's saying out loud: at both DeepSeek and Z.ai, the Flash tier ships with a 1,310,720-token context window while the flagship tops out at 1,048,576. The cheap model sees 25% more tokens than the expensive one. Whatever memory/throughput trick makes Flash fast is also buying it a bigger window. Pricing power hasn't just inverted — the spec sheet has too.
Implication: stop benchmark-shopping the flagship tiers. If Flash leapfrogging Pro is now a release cadence rather than an accident, the flagship exists mainly as a training-target and price anchor. Watch whether Anthropic and OpenAI respond with price cuts — or with anti-distillation hardening, which a few in today's HN thread expect. The open-weights labs have already voted.