Leadership1 publisher3 min readPublished
An Airbus SVP would split the robot stack into three models under independent safety controls
Greg Ombach, a senior vice president at Airbus, writes that vision-language-action models force a choice between capability and edge deployment, and that industrial buyers should validate one operating domain before expanding.
The Board Room · Leadership desk

What happened
- Airbus senior vice president Greg Ombach writes that larger vision-language-action models buy broader capability at the price of robot data and computing that make edge deployment difficult or impossible.
- He puts production automation at 80% to 90% in an automotive battery business he led and 20% to 30% in his aerospace experience, a spread of 50 to 70 percentage points.
- Public datasets hold more than 1 million real robot trajectories and cover only a fraction of industrial work, and physical interaction data is expensive, the column says.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- constraint If the transfer limitation holds, every new site is its own validation project, and a pilot's payback cannot be spread across a planned fleet of sites.
- decision A buyer specifying a cell this quarter has to decide whether the safety controls are contractually separable from the perception model, because that is what allows the model to be updated later without redoing the safety case.
- cost Generating the interaction data falls on the site owner, since public trajectory sets do not cover the work and simulated contact does not match it.
- exposure The 20% to 30% aerospace figure is one executive's account of his own career in a contributed column, so a business case that leans on it is leaning on testimony.
Where does the safety case live? That is the design question underneath the column. Ombach's stack has to run within independent safety controls, and when reality differs from the prediction it has to correct the action, request assistance or stop safely [6]. If the stop behaviour does not depend on the perception model being right, the model can be replaced without reopening the safety argument for the cell.
The split also changes what runs on every cycle. Many VLAs turn camera frames into visual tokens derived from pixels, then infer three-dimensional structure while deciding how to act [15]. His alternative gives object identity and task intent to a semantic layer, and position, orientation and velocity to a spatial model. A smaller task-focused model selects and adjusts movements without processing the full visual stream for every decision [5]. The geometry-first features he wants extracted are surfaces, edges, shape, position, orientation and movement, designed to stay stable across lighting and viewpoints [16].
Ombach draws the numbers from his own record. "I led an automotive battery business where we achieved 80% to 90% production automation," he wrote [7]. "In my aerospace experience, automation is closer to 20% to 30%," he wrote [8]. The spread is 50 to 70 percentage points [18]. The reason he gives is the product: millions of parts, much lower volumes, designs that prioritise aircraft performance, and changes that may require additional testing, documentation and approval [9].
A buyer can test the transfer claim. He writes that smaller local models cannot yet move learned skills across robots, tasks and sites without substantial retraining and validation [4]. Take a pilot cell's model to a second site and count the retraining and revalidation hours. If the second install costs close to the first, the multi-site payback in the business case rests on an assumption that has now been measured.
The cost of the data lands on whoever owns the site. Public datasets hold more than 1 million real robot trajectories and cover only a fraction of industrial work, and physical interaction data is expensive [14]. Simulation differs from reality in friction, mass and contact [20]. So the requirement he sets is a model that learns quickly from limited demonstrations and runs locally with low latency [19].
This is a contributed Forbes Tech Council column, written by a sitting executive, and it describes requirements for an architecture it does not show in service [c2, c19]. The column reports no deployment and no measured results. The sequencing advice stands on its own: define the operating domain, prove performance under real conditions, and expand only when more variability can be managed [13]. He takes the pattern from autonomous driving, where highway functions arrived before broader autonomy and commercial services began in mapped areas with speed limitations before expanding location by location [11]. General robotics is harder still, he holds, because it has no common rule set [12].
For a buyer deciding this quarter, that puts pilot scope ahead of model choice. His near-term band is tasks too variable for conventional automation but defined enough to validate economically and safely. Less structured industrial settings come next, and healthcare and consumer robotics only once cost, reliability, privacy, safety and security reach the required standards [10].
What to watch
- Whether Airbus or another manufacturer reports a fielded geometry-first cell with retraining or cycle-time figures attached.
- Whether robotics vendors begin quoting an independent safety layer as a separable line item.
- Whether public robot trajectory data grows past the 1 million count and starts covering industrial tasks.