BuildNot yet confirmed elsewhere1 publisher3 min readPublished
China's urban pilot-assist stacks give up on enumerating the intersection
Rule libraries scored on pre-mapped routes and asked for a takeover everywhere else. The pivot to one-stage end-to-end buys human-feeling control by deleting the per-module debug surface.
The Engineer · Build desk

What happened
- Early urban pilot-assist stacks in China tried to cover city driving with thousands of hand-written if-then-else rules for specific traffic situations.
- The source's account of Chinese city traffic, from wrong-way scooters to mid-block pedestrians, is a long tail no rule library can fully enumerate.
- The first end-to-end systems in production split the pipeline: sensors to an intermediate representation, then representation to trajectory.
Why it matters
- constraint Headcount cannot buy coverage here. Rules accumulate one scenario at a time while the road population producing scenarios is unconstrained by them, so the gap does not close with budget.
- decision The architecture choice is now a pricing decision: a stack whose failure mode is handing control back cannot be sold into segments where the buyer will not tolerate that.
- contradiction Two-stage was adopted partly because each module could be inspected; the fix for its lossy boundary removes those inspection points, and the source presents both moves as progress.
- exposure With no intermediate outputs, a bad intervention has no module to attribute it to, which puts the whole burden of explanation on fleet behaviour evidence.
A rule library is written by a fixed number of engineers. The situations it has to cover are produced by everybody else on the road, and in the Chinese city those authors are not working from the same document: scooters going the wrong way, pedestrians crossing mid-block, delivery riders weaving between cars, congested-intersection chicken that nobody is taught [5]. Each if-then-else statement buys coverage of one situation [4]. The population generating situations is not bounded by the rule set, so a bigger team writes more rules and still does not close the set. That is the difference between a hard problem and a structurally wrong method.
The failure mode described by one early test team is the diagnostic: the system scores beautifully on pre-mapped routes, then hesitates and asks the driver to take over the moment it meets something unrecorded [6]. A takeover request is not a slightly worse drive. It is the feature switching itself off in exactly the conditions the buyer was sold on.
Which is why the penetration figure carries more weight than the architecture argument. Urban NOA reached about 15.1 percent in China in 2025 and stayed concentrated in cars above RMB 200,000 to 300,000 [15]. The other 84.9 percent of the market did not have it [17], and the source's own reading is that the one-trick behaviour is the reason, with no route down-market until the long tail is handled [16]. Rule enumeration priced itself into a segment where the driver tolerates the handback.
The first production answer, two-stage end-to-end, moved the defect rather than removing it. Stage one compresses sensor data into a semantic map, drivable regions and detected objects; stage two turns that into a trajectory [8]. The module choosing what to keep does not know what the planner will need, and the boundary discards signal by construction [10]. The result was a car that stopped being rigid and started being clumsy: hard braking, hesitant lane changes, jerky longitudinal control [11].
One-stage removes the boundary, and with it the thing that made two-stage shippable. The clean intermediate products were what let engineers monitor and debug perception and planning independently [9]. A single network from sensor input to trajectory, with no manually defined intermediate module [13], has no such outputs to inspect. The validation burden moves off the unit test and onto fleet behaviour.
On what one-stage buys, the source is qualitative. It describes smooth following, early throttle lift and coordinated steering, the internalised vehicle sense of a human driver [14]. There is no takeover rate, no intervention distance, no comparison on the same routes [12]. So the consensus among leading players over the past year [13] currently rests on how the car feels, argued against a rules baseline whose weakness was measured in the same currency. That is a reasonable trade to make and an unreasonable one to verify, and the vendors who go first will be the ones explaining a bad intervention without a module to point at.
What to watch
- Whether any vendor publishes takeover or intervention rates for a one-stage stack against its own two-stage predecessor on the same city routes.
- Whether urban NOA appears below the RMB 200,000 band, and what the 2026 penetration number does off a 15.1% base.
- How engineers localise a bad one-stage intervention once the inspectable intermediate outputs are gone.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence28
- Adoption52
- Hype gap+38
- Incentives
- Insufficient
- Confidence34
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
China's intelligent driving is moving from highway Navigate on Autopilot (NOA) into urban NOA.
- [2]
Highway NOA was relatively simple to crack: the road is closed, the geometry consistent, the actors mostly cars, and a mature rule-based stack can deliver a comfortable product.
- [3]
Urban NOA must handle traffic lights, unprotected turns, pedestrians, e-bikes and food-delivery scooters running red lights, and complexity grows exponentially.
- [4]
The earliest urban NOA architectures enumerated traffic situations and encoded them as thousands of hand-written if-then-else statements, covering cases such as when to move off after a green light, how much to slow when cut off, and how to plan an unprotected left turn.
- [5]
Chinese urban road users largely do not follow the rules: electric scooters drive the wrong way, pedestrians cross mid-block, food-delivery riders weave between cars, and drivers play chicken in congested intersections. These are long-tail scenarios that no rule library can fully enumerate.
- [6]
An early city NOA test team said of their own system: "It feels like exam cramming - it scores beautifully on the routes we pre-mapped, and the moment it hits an unrecorded scenario, it hesitates, behaves awkwardly, and then asks the driver to take over."
- [7]
To escape hand-written rules the industry pivoted to end-to-end learning: a deep neural network replaces the rule stack, takes sensor input and directly outputs driving commands, trained on large volumes of human driving data.
- [8]
The first wave of end-to-end systems in production was almost always two-stage: the first stage turned raw sensor data into an intermediate representation (semantic map, drivable regions, detected objects), and the second stage turned that into a trajectory or control command.
- [9]
Two-stage end-to-end shipped faster and more explainably, because clean intermediate products let engineers monitor and debug perception and planning independently.
- [10]
The fatal flaw of two-stage was information loss: the intermediate representation chosen by the perception module is not guaranteed to contain every detail the planning module needs, and the module boundary is by construction a bottleneck that throws away signal.
- [11]
Two-stage improved user experience over pure rules but the car still drove stiffly, with hard braking, hesitant lane changes and jerky longitudinal control; it solved the rigid problem without solving the clumsy one.
- [12]
The source's case for one-stage end-to-end is qualitative: it offers no takeover rate, intervention distance or like-for-like route comparison against two-stage or rule-based stacks.
- [13]
Over the past year one-stage end-to-end has become the technical consensus among leading players: one neural network from sensor input to driving trajectory, with no manually defined intermediate module, optimised end-to-end against a safety objective.
- [14]
The most visible feature of one-stage end-to-end is how human it drives: the model internalises a vehicle sense, follows smoothly, anticipates the lead car, lifts off the throttle early and coordinates steering and throttle without jerky wheel input.
- [15]
Urban NOA penetration in China reached about 15.1% in 2025 and remained concentrated in vehicles priced above RMB 200,000 to 300,000.
- [16]
The source attributes that limited penetration to the one-trick ('偏科') experience, and states that until the system can handle the long tail it cannot move down-market.
- [17]
At about 15.1% urban NOA penetration in 2025, roughly 84.9% of the market did not have urban NOA.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toThe Evolution of China's Urban Pilot Assist: From "Exam Cramming" to One-Stage End-to-End
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.