Skip to content

Invest1 publisher2 min readPublished

RoboCurve times Astra's zero-shot picks at two and a half minutes each

RoboCurve scored 19 successes in 20 physical attempts on a pair of I2RT YAM arms, at roughly 2.1k output tokens a run. The humanoid pickup described in the same account happened in a simulator.

The Investor · Invest desk

Illustration accompanying RoboCurve times Astra's zero-shot picks at two and a half minutes each

What happened

  • cryptobriefing.com reports that third-party evaluator RoboCurve scored OpenAI's GPT-6 Astra at 19 successes in 20 physical pick-and-place attempts on dual I2RT YAM bimanual arms, a 95% rate.
  • Each of those runs consumed roughly 2.1k tokens of output and finished in approximately 2.5 minutes.
  • In RoboLab simulation benchmarks the same model completed 49 of 50 single-arm assignments, a 98% success rate.
  • The evaluations flagged persistent limitations on high-precision and complex bimanual manipulation.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint Two and a half minutes per pick allows 24 cycles an hour from one arm pair, a rate that suits low-volume, high-variety handling and sits far below line-rate packing.
  • decision An integrator can now weigh a per-task data collection and fine-tuning budget against a per-run token bill plus the staff time to clear roughly one failed pick in twenty.
  • contradiction The publisher credits Astra with controlling a full humanoid while its own text puts the physical 95% on two arms, so anyone budgeting off the headline is pricing hardware the evaluation did not measure.
  • exposure One evaluator's twenty attempts is the entire physical record behind the figure. Any capex case built on it is exposed to the next, larger test.

Two and a half minutes a cycle at 2.1k output tokens a run comes to roughly 50,400 output tokens an hour of arm time [3][11]. Run that arm pair around the clock and it emits about 1.21 million output tokens a day [18]. The account has the token count but no price per token, so the hourly inference bill cannot be worked out from it [17].

The physical evidence base is twenty attempts [2]. Eighteen successes out of twenty is 90% [12]. One more miss would have produced a much duller figure. cryptobriefing.com attributes the three points between the simulator and the bench to unexpected lighting and imperfect surfaces [15].

The humanoid item in the account is a zero-shot cola-bottle pickup in simulated control, translated from human demonstrations using camera input only [6]. The article's headline says Astra demonstrated zero-shot pick-and-place on a full humanoid robot [16]. Its text says comprehensive documentation of performance on full humanoids had yet to be revealed, with explorations still on arms and simulators [9]. The 95% belongs to the arms.

What a per-task pipeline buys is data. According to cryptobriefing.com, Astra used none of the usual kind: no privileged access to the robot's internal state, no specialized robotic datasets, just camera views and proprioception fed into actions [5]. The new spending line is tokens per pick plus somebody to clear the failures. A one-in-twenty miss rate at this cadence works out to 1.2 interventions an hour per arm pair [13].

RoboCurve also ran the predecessor Claude Fable variants and reports Astra ahead of them on speed, token efficiency and reliability, particularly on gross motor tasks [14].

In my view the binding number here is the two and a half minutes. The 95% is not. A dedicated pick cell is cheap per pick once it is running and expensive to set up for each new object. A general model that needs no dataset per object competes against that setup cost, and at this cadence setup cost is the only place it competes. The counter-thesis is that cycle time and token count per task fall faster than robot hardware costs do, and this evaluation already shows a generation-over-generation gain on both [14]. What would change my view is a physical sample in the hundreds holding at the same rate, or a published cycle time under a minute.

What to watch

  • Physical full-humanoid results from RoboCurve or a second evaluator, with the attempt count stated.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories