Product1 distinct publisher3 min readPublished
The figures are self-reported and simulated, but the metric choice is what travels: Sun Shengzhi's team scored its intent classifier on lethal weapons use under the command system it replaces, and never on speed.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
A watch officer has a contact the system rates 72% likely to be ordinary transit. Then it becomes three contacts fanning into a wedge on separate bearings, the same system moves coordinated probing to 68%, and it tells him to use the water cannon [6]. The output is a recommendation with a named effector attached, arriving at the second when a person is choosing between doing nothing and doing something that cannot be taken back.
The number that will travel is 0.3% going to zero [1]. It is also the softest number in the work. A 0.3% base rate is roughly three events per thousand runs [1], and the reported account carries no trial count [2], so a reader cannot tell whether the behaviour was designed out or the rare event simply did not come up. The accuracy comparison carries more weight: 89.6% against 68.5% is a gap of 21.1 points [5], which means misread intent falling from 31.5% of cases to 10.4%, about two thirds of the errors gone [3]. Separation violations going from 1.2% to 0.8% [5] is a third off [4], and it is the least promotable figure in the set, which is why it is worth reading closely. It suggests the model is changing what the hull does rather than only what the display says.
Teams buying decision support usually tell themselves they are buying time: faster classification, and the operator seeing it sooner. Sun Shengzhi's group, in a paper published this month in Command Control & Simulation, is claiming something else [4]. It claims a lower rate of the worst outcome, scored against the rule-based command process that produced the 0.3% in the first place [2], with the water cannon presented as the successful response rather than a fallback [12]. Speed is absent from the claim.
Which changes what a hard question looks like. Once a worst-outcome rate exists for the incumbent process, "our model is 89.6% accurate" stops being a sufficient answer, because accuracy says nothing about which errors are recoverable. A misread that fires a water cannon costs a diplomatic complaint. A misread one rung up costs something no review board can walk back.
The deception run is the part that would show up as behaviour on a real bridge. Fishing boats feign a retreat, other vessels accelerate in from another bearing, and the system holds position instead of reacting to the withdrawal [7]. That is a recommendation to do less, which is the hardest output for a tool to produce and the easiest for a commander to overrule. The authors say human commanders stay in the decision [8], so what they have measured is what the person on watch gets nudged toward.
Two axes sort this class of tool from the deck that sells it. One: is the baseline the process you run today, or one the vendor picked? Two: is the headline metric the rate of the outcome you cannot reverse, with a stated denominator, or accuracy and latency? This paper sits in the incumbent-baseline, worst-outcome cell, with the denominator missing. Everything in the other three cells is a demo with error bars you were not shown.
Ranked by verification strength, evidence, and original report placement.
A new AI system developed by Chinese researchers reduced the simulated risk of a Chinese coastguard vessel using weapons during disputed-water encounters from 0.3% to zero.
Under the existing rule-based command system, the simulated coastguard vessel used lethal weapons in 0.3% of scenarios.
In simulations, the AI correctly identified vessel intentions 89.6% of the time, compared with 68.5% for a conventional rule-based system.
The research was conducted by a team led by Sun Shengzhi, a professor at the China Coast Guard Academy, and published earlier this month in the Chinese journal Command Control & Simulation.
The AI also reduced other rule violations, such as vessels moving too close to one another, from 1.2% to 0.8%.
In one simulation the AI initially assigned a 72% probability that a nearby vessel was simply passing through; when three vessels formed a wedge and approached from different directions, it raised the probability of coordinated probing to 68% and ordered the use of a water cannon rather than escalating immediately.
Distinct publishers with included, body-backed reporting in this cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
product
Two Chinese booster recoveries, one state and one commercial: reflight is still the untested part1 distinct publisher
invest
SMIC's binding constraint is floor space, and the AI crunch has reached power chips1 distinct publisher
product
Southeast Asia's data centre pipeline is 173 announcements chasing 1,435 MW1 distinct publisher
invest
Singapore's deepfaked cabinet cost one victim S$4.9m, and the NDA did the real work1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One relay, paper unread
Every number in this story — 89.6, 68.5, 0.3, zero — reaches us through three hands: a Chinese-language paper in Command Control & Simulation, the South China Morning Post's reading of it, and Interesting Engineering's reading of that. Nobody in our coverage has the paper. The percentages arrive with no run count, no description of the rule-based system they beat, and no third party who has looked at the method.
Nothing has touched water
What exists is a scored war game and a stated intention to keep human commanders in the loop. There is no sea trial, no vessel fitted, no unit using it, no procurement and no date — and the reporting does not claim otherwise. The only observable event is publication.
The metric flatters the system
Look at what the scoreboard measures. The team graded its intent classifier on how often lethal weapons come out, against the command system it wants to replace — and nothing on the sheet penalises hesitation, slow decisions, or a probe waved through as harmless transit. A cautious model wins this contest by construction. The write-up then compounds it: the headline hands coast guard academy work to the Chinese Navy and leads with 0.3%, the old system's failure rate, while the promise that this reduces real-world accidental escalation has been tested nowhere but inside the simulator.
Graded by its own builders
The graders here are the graded. Sun Shengzhi's team at the China Coast Guard Academy chose the scenario, the baseline, the metric and the venue, and reports that its own system scored perfectly on the measure it selected. The finding also happens to be diplomatically convenient: a Chinese coast guard result showing machine-assisted restraint in disputed waters. None of that makes the numbers wrong; it does mean nobody with a reason to look for problems has looked.
Low, and untested from outside
One thread of reporting, simulated numbers reported by the people who built the thing, a zero without a denominator, no replication. Notably there is also nothing contradicting it — which in a single-source story is not reassurance, just silence.