Skip to content

Invest1 publisher3 min readPublished

Robot makers buy the inference before they sell the robot

Where a robot's inference runs decides what the machine costs to build. SemiAnalysis puts the industry's upfront compute bill in the trillions at a billion robots, paid by manufacturers on every unit before it does any work.

The Investor · Invest desk

Photograph accompanying Robot makers buy the inference before they sell the robot
Photo: figure.ai

What happened

  • SemiAnalysis names two constraints that invert the LLM order of operations for robots: a real-time control loop that cannot miss a deadline, and compute the manufacturer builds and pays for on every unit.
  • Generalist robot models sit in the billions of parameters, with Physical Intelligence's pi-0 near 3 billion, ByteDance's GR-3 at 4 billion, Generalist around 10 billion and NVIDIA's DreamZero at 14 billion.
  • Frontier robot models already outgrow the robot they drive: pi-0.7 runs on an off-robot H100, and DreamZero needs two GB200 GPUs off the robot to run in real time.
  • DreamZero is a 14-billion-parameter world action model built on a video-diffusion backbone, which is what makes its real-time compute requirement so heavy.
  • SemiAnalysis expects a mix, with some robots running cognition entirely onboard and others offloading to datacenter GPUs that pool inference across a whole fleet.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • cost An LLM company's serving cost arrives with the query, on a screen the customer already owns. A robot maker's arrives with the unit, so selling more machines consumes cash before it returns any.
  • constraint A robotics lab cannot spend its way past the deadline the way an LLM lab spends on training, because capability is capped at what runs in real time on hardware cheap enough to ship in volume.
  • decision Inference placement has to be settled before the bill of materials is fixed. Onboard silicon is a cost repeated on every unit, while datacenter GPUs are shared capital serving many robots.
  • contradiction NVIDIA's two models point opposite ways within months, so the evidence here does not settle whether inference runs onboard, and SemiAnalysis says the field has yet to converge.

Take the top of the range SemiAnalysis gives and divide it. The newsletter puts the upfront compute cost at billions, and in the trillions if the world ever gets to a billion robots [3]. A trillion across a billion units is a thousand a unit [18]. The manufacturer builds and pays for that on every machine (the newsletter does not state a currency). An LLM's inference runs behind a screen the user provides [4].

The other constraint is time, and the deadline holds at any budget. A robot runs real-time control loops and cannot miss a deadline, and when it is slow the world changes around it, so the action is obsolete by the time it has to be made [2].

Size therefore gets picked differently. LLM labs choose parameter counts by trading quality against a training budget and an inference cost. Robotics labs choose by what their data supports and what fits on a Jetson or an H100 inside a latency budget, according to SemiAnalysis [8]. The counts still tell you something. Fourteen billion parameters against a frontier LLM at a trillion is a factor of about 71 [19]. Inside one lab's line-up, the step from Physical Intelligence's π0 to π0.7 is 3 billion to 5 billion, or 67 percent [21]. SemiAnalysis's own caveat is that these numbers show where the constraints sit today, and where sizes will settle is still open [22].

Right now they sit off the robot. Two GB200s dedicated to one machine would dominate its bill of materials. Pooled across a fleet, the same pair serves many, and pooling inference is one of the two advantages SemiAnalysis claims for datacenter compute, the other being escape from the robot's compute and power budget [16]. Pooling divides the bill only when the fleet wants its GPUs at different moments. The link brings latency and jitter of its own [14].

I think placement is the decision that governs who can build robots in volume, and the strongest argument against that is in the same source. RoboTTT's 3 billion parameters are about a fifth of DreamZero's 14 billion [20], and it gets minutes of usable context by continually updating its own weights at test time instead of generating videos of the future [11]. If that direction holds as tasks generalise, the off-robot GPU is a stopgap and the network question fades. SemiAnalysis wrote that "our view is that a cascade of approaches is inevitable" [15], and says nobody has settled the hardware, the models, or the economics [17].

What decides it is data. There is no internet-scale corpus of robot experience, and real-world interaction data is slow and expensive to generate; SemiAnalysis credits that gap for the rise of embodied human data collection and world models [12]. Early results from Generalist and Dyna suggest robotics scales the way LLMs did [13].

What to watch

  • Whether NVIDIA's test time training line scales to generalist tasks or stays a 3-billion-parameter demonstration.
  • A robot maker disclosing per-unit compute cost inside a bill of materials, putting a real figure against SemiAnalysis's trillions bracket.
  • Whether pi-0.7's off-robot H100 requirement moves onboard in a later version of the model.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories