Total unique visitors
Browse by category Chatbots Image Generation Video Generation Audio & Voice Coding Writing Productivity Research AI Agents Free Tier Table
Home page 問問貓說 AI

Ask CatAI Tool SummaryChatGPT

GPT-6 Astra's 19/20 Robot-Arm Result Went Viral — The Author Listed Four Caveats Nobody Quoted

Article last updated:2026-09-07

On 2026-09-04 a report titled “GPT-6 Astra on robotic manipulation” hit 229 points on Hacker News. The headline number is easy to repost: OpenAI’s new flagship placed a block in a bowl 19 times out of 20, while Claude Fable 5.1 managed 8.

We read the original report and copied out the full table — completion rates, cost, runtime — along with the four limitations the author himself wrote into the piece. Verified 2026-09-07.

1. How the experiment actually ran

The report is a follow-up to Robocurve’s earlier Claude Fable 5 vs Fable 5.1 comparison, so the hardware, tasks and agent policy are carried over unchanged; only the new model was added.

  • Hardware: bimanual I2RT YAM arms, 6 degrees of freedom per arm, parallel-jaw grippers.
  • Agent policy: the same Inspect Robots agent policy as the previous round.
  • Task one: pick up the red block from the table and place it inside the bowl.
  • Task two: pick up the round blue puzzle piece by the knob at its centre and place it into the matching circular groove.
  • Scale: three models, two tasks, 20 trials each — 120 trials in total.
  • Scoring: a human grader records the highest stage reached (0 no purposeful approach, 1 contact, 2 lifted clear of the table, 3 positioned above the deposit point, 4 placed), so failed runs still record how far they got.

2. The full numbers, not just the 19/20

TaskModelMean stageCompletionsRateOutput tokens/runEst. cost/runMinutes/run
Block into bowlFable 51.301/205%19.2kUS$2.698.2
Block into bowlFable 5.12.408/2040%12.9kUS$2.126.8
Block into bowlGPT-6 Astra3.9519/2095%2.1kUS$0.942.5
Puzzle into grooveFable 51.500/200%16.3kUS$2.637.9
Puzzle into grooveFable 5.12.352/2010%10.5kUS$2.185.9
Puzzle into grooveGPT-6 Astra2.002/2010%2.7kUS$1.363.4

Only the first block travelled. The second task is the one worth reading: every model failed it. Astra completed 2 of 20; so did Fable 5.1. The report says plainly that Astra reaches the groove and stalls at the same final step Fable does.

So what the data actually shows is narrower than the headline: on simple pick-and-place, Astra pulls clearly ahead; on precision insertion, nothing here works yet.

3. The column that matters more than the completion rate

The most informative column is the one nobody quoted — output tokens per run:

  • Block into bowl: Astra 2.1k, Fable 5.1 12.9k, Fable 5 19.2k — a 6× to 9× gap.
  • Runtime tracks it: 2.5 minutes against 6.8 and 8.2.

That is where the report’s “2.3× cheaper, 2.4× higher completion rate” summary comes from. For anyone costing out a real workflow, this column is the practical one: the model that says less is not just faster, it bills less.

4. The four caveats the author states outright

The original report contains a limitations section that almost no repost carried:

  1. Astra’s trials ran two days after the Fable trials, and were not interleaved. Environmental drift between batches cannot be ruled out.
  2. Grading was operator-judged with the model known, which the author describes as leaving scores “open to unconscious bias”.
  3. Costs use list price. The author adds that if anything, Astra’s cost is overstated.
  4. All models ran at medium reasoning effort only. Higher-effort settings were not tested.

Points 1 and 2 are exactly what a formal benchmark would be asked to fix — interleaved runs and blind grading. Disclosing them makes this report more honest than most vendor-adjacent evaluations. It also means the 19/20 should not be quoted as settled.

5. What this means if you are not building robots

Three things:

  • This is not an official OpenAI robotics benchmark. It is one third party’s rig and agent policy.
  • “Robot arms” does not mean a product you can buy. What was tested is a model driving existing arms through an agent policy.
  • The cost and runtime columns are the transferable part. If you are weighing a model switch, “6× fewer output tokens on the same task” converts to your invoice far more directly than a completion rate does.

For who can currently access GPT-6 Astra and how its API billing threshold works, see our read of the official model page: GPT-6 Astra is live but subscribers still can’t reach it, and the API doubles past 272K.


Sources: the original Robocurve report, published 2026-09-04, read directly, alongside its Hacker News discussion. Verified 2026-09-07. This is a third-party test rather than an official OpenAI benchmark, and the author discloses four experimental limitations — quote it with those attached. Plan and model pricing is summarised on our ChatGPT tool page.

What Amo and Pimi think

AMO Amo Finding faults
皮米,你等一下要誇免費版功能全對吧?我先講一個更陰的:GPT-5.3 用完約 10 則就自動偷偷降級成小模型,回答品質明顯變差,但介面完全不會告訴你已經換人回答了。額度也不是每天午夜歸零,是滾動式 5 小時窗,用完得乖乖等。更少人知道的是——你在設定裡關掉個人化廣告,額度會被砍到每 3 小時只剩 5 則。方案命名也亂,Pro 和 Pro Max 官網文案其實只寫「From US$100/month」,第三方說有兩檔不同價,我們查到現在都沒能一手確認。
PIMI Pimi Advantages
你講的降級是真的,但整體來看免費版還是功能最完整的一個——對話、搜尋、生圖、語音、檔案上傳全都給,別家通常挑一兩樣就開始收費了。額度用完也不是斷線,降級後照樣能用,臨時問一句話很夠。中文支援是完整等級,台灣用信用卡直接付。而且免費版產出可商用、沒有浮水印——阿莫你剛剛講降級講得很兇,但這點很多生圖工具連付費版都做不到。
So, do you need to pay or not?

If you only ask a question or two occasionally, the free plan is enough. If you use it daily for work and can't stand being downgraded mid-conversation, the US$20 Plus plan is the safest entry-level choice on the whole site. The Taiwan official site now prices in NT dollars — Plus is NT$690/month, the same as the App Store in-app purchase price, so it doesn't matter which one you subscribe through. If you want to save a bit, go with the App Store's annual billing at NT$6,990 (about NT$583/month).

Let's take a look at these

Go to the official website

Affiliate Links Notice