Z.ai GLM 5.3 Flash hits OpenRouter — 1M context at $0.075/M tokens

Z.ai's GLM 5.3 Flash went live on OpenRouter today, bringing a native multimodal model (text+image+video→text) with a 1M-token context window and pricing that undercuts almost everything at its capability tier.

What changed: The model clocks in at $0.075/M input tokens and $0.25/M output tokens — roughly 5× cheaper than the full GLM 5.3 on both axes. It supports text, image, and video inputs with text output, and uses a hybrid sparse-and-linear attention architecture designed for long-context coherence. The full model ID on OpenRouter is z-ai/glm-5.3-flash.

Why a builder cares: At this price point with a 1M context window and multimodal input, GLM 5.3 Flash makes long-horizon agent tasks and multi-turn codegen economically viable at scale. It slots into the same niche as Gemini 2.5 Flash but competes on price — and for teams running heavy context workloads, the cost difference adds up fast.