GPT-6 Astra on Robot Arms: 19/20, $0.94 Per Run

Robocurve put GPT-6 Astra on a pair of bimanual I2RT YAM robot arms under the same Inspect Robots agent policy it already used for Claude Fable 5.1. Same two tasks, 20 trials each, human-graded on a five-stage rubric. The scoreboard: 19 of 20 block-into-bowl placements against Fable 5.1's 8 of 20 — in 2.5 minutes per run instead of 6.8, at $0.94 per run instead of $2.12. Then both models hit the identical wall on the precision task: 2 of 20 each.

What Robocurve Measured

Six-DoF arms controlled with absolute end-effector poses, three camera views plus proprioception, medium reasoning effort, a 20-LLM-call budget, 120 counted trials total. The part everyone scrolled past is the token column: Astra finished a bowl run on ~2.1k output tokens versus Fable 5.1's 12.9k — 84% fewer. Divided by completions, that's roughly $0.99 per successful placement versus Fable 5.1's $5.30. Efficiency compounds accuracy here: the model that knows what to do stops narrating its way there.

Why It Matters

The puzzle piece — lift by a center knob, seat it in a circular groove — is the interesting failure. Astra reaches the groove and stalls at the same final step Fable does; mean stage 2.00 versus 2.35. More intelligence bought nothing on last-inch contact. Robotics splits into two problems: coarse pick-and-place, where frontier quality now completes at under a dollar a task, and force-fit insertion, where no frontier model has moved the number. That gap is why the laundry-folding livestreams still film in controlled rooms.

Receipts cut both ways, so read the caveats: the bowl comparison ran on a different rig than the Fable trials, grading was operator-judged with models known, trials weren't interleaved, and Astra's automatic prompt caching wasn't discounted — its cost edge is, if anything, understated.

Verdict

The step change is real but narrow. Astra made "reasoning model drives a robot arm through a pick-and-place" a solved-and-cheap category, and simultaneously proved the dexterity frontier hasn't budged. Watch who attacks stage-4 contact next — that's the only number left on this board.