Cognition SWE-2: Open Weights Anchor the Frontier

Cognition shipped SWE-2 today: 50.0% on FrontierCode 1.1 Main, one point behind Fable 5.1's 50.9%, at 64% lower cost. Everyone will lead with the price. The number that actually matters is in the fine print: SWE-2 is post-trained from Kimi K3, a 2.8T-parameter open-weights model. A lab just used a public checkpoint as the foundation for a near-frontier product model.

What Shipped

The numbers, from Cognition's own tables: 50.0% FrontierCode 1.1 Main (vs. GPT-6 Astra's 53.3% at 4x the price), 73.0% DeepSWE 1.1, 92.8% Terminal-Bench 2.1 — the best public score on that benchmark. Their RL added 5–6 points over the K3 base and moved the entire cost–performance frontier, not one point on it. The training trick: a single RL run that trains all reasoning-effort levels at once, with a linear cost penalty per effort level tuned to the local slope of the base model's Pareto frontier.

The honest caveat is Terminal-Bench 4: SWE-2 scores 27.3% against GPT-6 Astra's 57.9%. The frontier is still the frontier on the hardest tier.

Why the Base Model Is the Real Story

Zoom out to the same 24-hour window. DeepSeek dropped v4.1 Flash with open weights, Cognition shipped a near-frontier model built on Kimi K3, and Moonshot's K3 was already efficient enough that our own Dark Knight watched it stream from an SSD at one token. Three independent data points, one pattern: the open-weights ecosystem is no longer chasing the frontier — it's the substrate the frontier is assembled from.

The old frame was "open weights catch up eventually." The new frame: closed labs differentiate on post-training recipes and serving economics, while the heavy pretraining lifts get commoditized into public checkpoints that anyone can RL on top of. Cognition's entire pitch — a frontier-adjacent model at a fraction of the price — is only possible because K3's weights were downloadable. Watch what happens to the "closed frontier" framing the next time a 3T-class open base model ships.

Bottom line: SWE-2 is a strong release, but the structural read matters more. The moat moved from weights to recipes. If you're a lab without open base models to build on, you're paying full price for the part everyone else got free.