Product1 publisher3 min readPublished
J-20's design institute says model hallucination is a requirements problem, not a productivity one
Engineers at AVIC Chengdu warn that language models invent radar ranges, payload figures and fatigue limits with full confidence, and that fabricated threat data lands upstream of design decisions.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Engineers at China's AVIC Chengdu Aircraft Research and Design Institute, the organization behind the J-20 and a key developer of the next-generation J-36 fighter program, warned that artificial intelligence could feed false data into military intelligence, potentially distorting the design and operation of advanced weapons.
- The warning came in a paper published on June 20 in Information Studies: Theory & Application, a journal operated by China's state-owned defense manufacturer Norinco.
- The concern centers on large language models, which can rapidly analyze vast amounts of information but also generate convincing answers that have no factual basis.
- For fighter aircraft programs, such mistakes could affect assessments of enemy radars, missile capabilities and electronic warfare systems, and engineers could then make decisions using intelligence that looks credible but fails basic technical checks.
- Military aircraft development begins with a detailed understanding of the threats an aircraft may face, requiring accurate information about enemy radar coverage and operating frequencies, missile engagement zones and the capabilities of electronic warfare systems.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
Engineers at China's AVIC Chengdu Aircraft Research and Design Institute, the organization behind the J-20 and a key developer of the next-generation J-36 fighter program, have published a warning that artificial intelligence can feed false data into military intelligence and distort the design and operation of advanced weapons [1]. The paper appeared on June 20 in Information Studies: Theory & Application, a journal operated by the Chinese state-owned defense manufacturer Norinco [2].
What makes this different from the usual complaint about chatbots making things up is where the error enters. Military aircraft development begins with a threat picture: enemy radar coverage and operating frequencies, missile engagement zones, and electronic warfare capability [5]. The positions of early-warning aircraft and ground-based air defenses shape requirements too, and those findings drive choices about stealth, sensors, maneuverability and combat range, so a flawed assessment can affect a program before the requirements are finalized [6]. A fabrication at that stage does not surface as a bad answer in a chat window. It surfaces as a specification.
The paper's author, Zhang Xianzhe, an engineer connected to the Chengdu design establishment, warns that large language models can invent highly specific military information: stealth fighter dimensions, payload figures, weapon capabilities, top speeds, combat radius [7]. A military intelligence system may report incorrect radar scan ranges or frequency bands while presenting them with complete confidence, convincing enough to escape immediate scrutiny [8]. That is the operative property. The output is wrong in the format of a correct answer, which is exactly the failure mode that review processes tuned to catch obvious nonsense will pass.
The design side of the risk is more concrete. According to Zhang, a model could recommend materials that exceed their actual fatigue limits or suggest maneuvers that violate aerodynamic constraints [9], and could produce mission plans that send an aircraft beyond its real operational range [10]. On the intelligence side, systems may generate bases, units or troop movements that never existed, and analysts feeding those outputs into simulations would then produce false assessments of an opponent's real combat capability [11]. Those errors are load-bearing in a way a wrong summary is not: they propagate into structures, envelopes and plans.
The institute is also blunt about the limits of the standard mitigation. A human analyst may hold final authority, but reviewing every AI-generated recommendation gets harder as the volume and speed of information rise [12]. Human-in-the-loop is a throughput claim disguised as a safety claim, and it degrades precisely when the system is doing the work it was bought to do. The report adds that a high-casualty strike earlier this year reportedly followed an AI system's identification of a target as high priority, alongside outdated intelligence [13]; no further identifying detail is given, so treat that as an unverified single-source assertion rather than a case study.
Zhang's proposed safeguards will look familiar to anyone who has shipped a retrieval system: ground the models in trusted military data, build searchable knowledge bases for verification, write clearer prompts to reduce ambiguity, and cross-check conclusions using models that challenge each other's outputs [14]. What the account does not contain is a residual error rate, an acceptance threshold, or evidence that adversarial cross-checking catches confident fabrication rather than laundering it.
Worth watching: whether any program turns this into a gate rather than guidance, meaning provenance tags on threat parameters, verification bases treated as configuration items, and sign-off that names the source of each number. The venue matters too. Guidance published in a Norinco journal is closer to procurement culture than to research commentary. The report frames the same exposure for the U.S. military as AI moves into intelligence analysis, weapons development and battlefield planning [15], where the equivalent question is whether any Western program office has written down what an unverifiable model output disqualifies.