GLM-5.3-Flash: 320B MoE, 18B Active, Approaches Opus 4.8
Twelve days after GLM 5.3, Zhipu AI (z.ai) is back with GLM-5.3-Flash — and the cadence alone is a signal worth watching.
GLM-5.3-Flash is a 320B-parameter MoE model with just 18B active parameters per token. It's the first natively multimodal model in the GLM-5 series (image-text-to-text), licensed MIT, with weights live on Hugging Face.
The numbers are what make this interesting:
- DeepSWE v1.1: 63.4 — competitive with models 10× its active cost
- TerminalBench 2.1: 84.3
- HLE (with tools): 55.3
- Pricing: ~$0.50/M output tokens via Novita, ~$0.50 via Baseten
- Speed: ~34 tok/s on z.ai's own hardware, ~139 tok/s on Baseten
On DeepSWE's leaderboard, GLM-5.3-Flash sits alongside gpt-5.6-luna, claude-sonnet-5, and gemini-3.5-flash. HN commenters report it "smashes deepseek-v4-flash and matches deepseek-v4-pro at a fraction of the cost." The model supports tool calling and structured output via a Jinja-based chat template.
The bigger story here is the trajectory. Look at the timeline from one HN watcher:
July 16: Kimi K3 — "China has caught up to Opus"
~4 weeks later: GLM 5.3 — same performance, ⅓ the parameters and cost
12 days later: GLM 5.3 Flash — almost same performance, half the active params again, ⅕ the price, on Chinese chips
That last part matters. This isn't running on H100s — Zhipu is serving inference on domestic silicon. The model has 62 shards of safetensors (321B parameters total, FP8 quant), natively multimodal, MIT licensed, available today. The gap between "frontier" and "cheap enough to not think about" is closing faster than the chart-watchers expected.
Verdict: GLM-5.3-Flash is the strongest "open-ish" model release this month. It won't beat Opus 5 or GPT-5.6 Luna Max in head-to-head benchmarks, but it doesn't need to — it lives in the cost-performance sweet spot that actually ships products. 943 HN upvotes and 474 comments in under 24 hours says the crowd agrees.