Mercury 2.5 Pricing on OpenRouter: 84% Cheaper Than Mercury 2
OpenRouter added Inception Mercury 2.5 (canonical slug inception/mercury-2.5-20260908) on Sept 8, and the list prices are a generational cut, not a bump. We pulled both generations' rates straight from OpenRouter's /api/v1/models endpoint.
What shipped
Mercury 2.5, the latest diffusion LLM, went live on OpenRouter with 260k context at $0.40 per million input tokens and $1.50 per million output tokens, with cache reads at $0.04/M. Its predecessor, Mercury 2, still lists at $2.50/M input, $7.50/M output, $0.25/M cache, on 128k context.
What changed
That's 84% cheaper input, 80% cheaper output, and 84% cheaper cache reads — plus a 2.03x context jump. Blended at a 3:1 input/output ratio, cost per million drops from $3.75 to $0.675, an 82% reduction in one release. For scale: 10M input + 2.5M output tokens now costs $7.75 instead of $43.75. Inception didn't just match the fast-model market (Qwen3-Next-Instruct sits at $0.90/M input, $1.10/M output on OpenRouter) — Mercury 2.5 undercuts it on input by 56% while charging a $0.40/M output premium for speed most samplers can't reach.
Why a builder cares
Mercury 2's price was the tax on diffusion latency. That tax is gone: high-throughput agents and autocomplete-class workloads can now route to Mercury 2.5 at commodity-input prices, and anything still pinned to inception/mercury-2 is burning 5x on output for no reason. OpenRouter also shipped two free Nex-N2.5 variants (262k context, $0/M both directions) the same day — the free tier and the cheap-fast tier moved in the same 24 hours. Trust stake: prices listed on OpenRouter are router rates; Inception's direct API may differ, so re-verify before committing budgets.