Build1 publisher3 min readPublished
A per-step 66 percent has to compound across the dozens of steps in Skild's S1 tasks
Skild says a single video demonstration teaches its S1 model an unseen ten-minute task with no weight updates. The accuracy figure published alongside that work is 66 percent per manipulation step.
The Engineer · Build desk

What happened
- Skild AI launched S1, a robot foundation model that takes a video demonstration as a prompt and executes the task without updating its weights or running task-specific post-training.
- Skild says S1 handles unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making and kit assembly, spanning dozens of manipulation steps.
- In Skild's own tests on new multistep tasks, S1 succeeded about 66 percent of the time at each step, against 9 percent for what the announcement calls a similar AI system.
- In one plant-potting test, the team went from recording the demonstration to autonomous execution on hardware in 11 minutes.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision An integrator buying on a per-step number has to run its own task-level acceptance trials before committing a cell, because nothing published says how often a ten-minute job finishes.
- cost The budget at risk is teleoperation data collection and the per-change training run; the sign-off testing after each layout change still has to be paid for by the plant.
- constraint With the 9 percent system unnamed, nobody outside Skild can reproduce the sevenfold gap, so it does not function as a procurement comparison against a specific alternative.
- capability If in-context task acquisition holds on hardware, a factory can add a task without owning a data pipeline or any model training capability at all.
Hand S1 a video and nothing inside the model changes. The recording goes in as a prompt, and the model reads intent, objects and sequence out of the frames, then maps them to actions for the robot in front of it, with no retraining and often for a task outside its pretraining set [3]. Skild calls this in-context learning [2]. The operator supplies no new dataset and triggers no training run [10].
The accuracy figure attached to that work is measured at each step [6]. Per-step success and task completion are different quantities. Assume steps are independent and one failure ends the attempt: 66 percent per step gives about 1.6 percent over ten steps and about 0.02 percent over twenty [1][2]. The tasks Skild describes run up to ten minutes and can span dozens of manipulation steps [4].
Those assumptions are the ones worth arguing about. Skild says the model adjusts when objects move and recovers from errors [5]. If a failed step can be retried in place until it lands, the multiplication does not apply, and per-step accuracy mostly sets cycle time instead. The announcement gives the per-step number without an end-to-end completion rate for the ten-minute tasks and without the trial count behind the 66 percent [5].
The comparison system is described only as a similar AI system [6]. At 9 percent per step, a five-step task finishes about six times in a million attempts [3], so step granularity is the only granularity at which that baseline is distinguishable from zero.
The figure a plant manager can act on is the data-collection one. Skild estimates that one short video can be as useful as roughly 380 hands-on training examples, and that a person collecting those manually could take 50 to 100 hours [7]. That works out to 8 to 16 minutes per example [4]. For the saving to transfer to a given line, the teleoperation pipeline there has to cost about the same per demonstration, and the new task has to need about 380 of them.
"Learning by experience, and not preprogramming, is the step change that has happened in robotics," said Deepak Pathak, cofounder and CEO of Skild AI [8]. The tests and the estimates are Skild's own, published on NVIDIA's blog, and Skild built S1 and ran the research on NVIDIA AI infrastructure as part of a collaboration covering synthetic data, training, simulation and deployment [14]. Cosmos world foundation models diversify the training data and turn video into structured descriptions, and Skild validates behaviour in Isaac Sim and Omniverse before hardware [15].
The production example in the post is the Foxconn work, where Skild, NVIDIA and Foxconn are putting Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems; in one demonstrated workflow a robot installs a busbar and limit block, fastens 16 screws and adapts to disturbances [9]. Skild says it reached a $100 million annual revenue run rate 10 months after its first commercial deployment [12], across more than 60 deployment partnerships spanning manufacturing, logistics, inspection, security and food preparation [13].
What to watch
- End-to-end completion rates and trial counts for the ten-minute tasks, which would settle whether per-step failures compound or get retried.
- Whether the 9 percent baseline is named, or whether anyone outside Skild reruns the multistep comparison.
- Cycle time and first-pass yield from the Foxconn Blackwell assembly cell with Skild Brain in the loop.