Ask Cat › AI Tool Summary › ChatGPT
GPT-6 Astra's 19/20 Robot-Arm Result Went Viral — The Author Listed Four Caveats Nobody Quoted
Article last updated:2026-09-07
On 2026-09-04 a report titled “GPT-6 Astra on robotic manipulation” hit 229 points on Hacker News. The headline number is easy to repost: OpenAI’s new flagship placed a block in a bowl 19 times out of 20, while Claude Fable 5.1 managed 8.
We read the original report and copied out the full table — completion rates, cost, runtime — along with the four limitations the author himself wrote into the piece. Verified 2026-09-07.
1. How the experiment actually ran
The report is a follow-up to Robocurve’s earlier Claude Fable 5 vs Fable 5.1 comparison, so the hardware, tasks and agent policy are carried over unchanged; only the new model was added.
- Hardware: bimanual I2RT YAM arms, 6 degrees of freedom per arm, parallel-jaw grippers.
- Agent policy: the same Inspect Robots agent policy as the previous round.
- Task one: pick up the red block from the table and place it inside the bowl.
- Task two: pick up the round blue puzzle piece by the knob at its centre and place it into the matching circular groove.
- Scale: three models, two tasks, 20 trials each — 120 trials in total.
- Scoring: a human grader records the highest stage reached (0 no purposeful approach, 1 contact, 2 lifted clear of the table, 3 positioned above the deposit point, 4 placed), so failed runs still record how far they got.
2. The full numbers, not just the 19/20
| Task | Model | Mean stage | Completions | Rate | Output tokens/run | Est. cost/run | Minutes/run |
|---|---|---|---|---|---|---|---|
| Block into bowl | Fable 5 | 1.30 | 1/20 | 5% | 19.2k | US$2.69 | 8.2 |
| Block into bowl | Fable 5.1 | 2.40 | 8/20 | 40% | 12.9k | US$2.12 | 6.8 |
| Block into bowl | GPT-6 Astra | 3.95 | 19/20 | 95% | 2.1k | US$0.94 | 2.5 |
| Puzzle into groove | Fable 5 | 1.50 | 0/20 | 0% | 16.3k | US$2.63 | 7.9 |
| Puzzle into groove | Fable 5.1 | 2.35 | 2/20 | 10% | 10.5k | US$2.18 | 5.9 |
| Puzzle into groove | GPT-6 Astra | 2.00 | 2/20 | 10% | 2.7k | US$1.36 | 3.4 |
Only the first block travelled. The second task is the one worth reading: every model failed it. Astra completed 2 of 20; so did Fable 5.1. The report says plainly that Astra reaches the groove and stalls at the same final step Fable does.
So what the data actually shows is narrower than the headline: on simple pick-and-place, Astra pulls clearly ahead; on precision insertion, nothing here works yet.
3. The column that matters more than the completion rate
The most informative column is the one nobody quoted — output tokens per run:
- Block into bowl: Astra 2.1k, Fable 5.1 12.9k, Fable 5 19.2k — a 6× to 9× gap.
- Runtime tracks it: 2.5 minutes against 6.8 and 8.2.
That is where the report’s “2.3× cheaper, 2.4× higher completion rate” summary comes from. For anyone costing out a real workflow, this column is the practical one: the model that says less is not just faster, it bills less.
4. The four caveats the author states outright
The original report contains a limitations section that almost no repost carried:
- Astra’s trials ran two days after the Fable trials, and were not interleaved. Environmental drift between batches cannot be ruled out.
- Grading was operator-judged with the model known, which the author describes as leaving scores “open to unconscious bias”.
- Costs use list price. The author adds that if anything, Astra’s cost is overstated.
- All models ran at medium reasoning effort only. Higher-effort settings were not tested.
Points 1 and 2 are exactly what a formal benchmark would be asked to fix — interleaved runs and blind grading. Disclosing them makes this report more honest than most vendor-adjacent evaluations. It also means the 19/20 should not be quoted as settled.
5. What this means if you are not building robots
Three things:
- This is not an official OpenAI robotics benchmark. It is one third party’s rig and agent policy.
- “Robot arms” does not mean a product you can buy. What was tested is a model driving existing arms through an agent policy.
- The cost and runtime columns are the transferable part. If you are weighing a model switch, “6× fewer output tokens on the same task” converts to your invoice far more directly than a completion rate does.
For who can currently access GPT-6 Astra and how its API billing threshold works, see our read of the official model page: GPT-6 Astra is live but subscribers still can’t reach it, and the API doubles past 272K.
Sources: the original Robocurve report, published 2026-09-04, read directly, alongside its Hacker News discussion. Verified 2026-09-07. This is a third-party test rather than an official OpenAI benchmark, and the author discloses four experimental limitations — quote it with those attached. Plan and model pricing is summarised on our ChatGPT tool page.
What Amo and Pimi think
If you only ask a question or two occasionally, the free plan is enough. If you use it daily for work and can't stand being downgraded mid-conversation, the US$20 Plus plan is the safest entry-level choice on the whole site. The Taiwan official site now prices in NT dollars — Plus is NT$690/month, the same as the App Store in-app purchase price, so it doesn't matter which one you subscribe through. If you want to save a bit, go with the App Store's annual billing at NT$6,990 (about NT$583/month).
Let's take a look at these
- ChatGPT Comprehensive Introduction: Pricing, Features, and Actual Limitations
- ChatGPT Is the free quota enough?
- ChatGPT Alternatives
- Comprehensive Free Quota List for All Tools

