Skip to content

Product4 publishers3 min readPublished

A formatting prompt carried an AI hallucination across US command channels

CNN reports that a Special Operations Command analyst used a chatbot to fuse open-source and classified signals intelligence, then used it again to turn the wrong answer into the summary that moved up the chain.

The Product Desk · Product desk

Illustration accompanying A formatting prompt carried an AI hallucination across US command channels

What happened

  • The Special Operations Command analyst then ran the tool a second time to format the erroneous finding into an official-looking summary, and that document was circulated across command channels.
  • CNN could not determine what the ship was actually carrying, or whether the analyst used a commercial chatbot or one designed for the US government.
  • Defense Secretary Pete Hegseth's January memo directed the department to become AI-first and to experiment with models from leading US AI companies.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Every team shipping a step that reformats a note into the house template now owns a labeling decision, because that step is where a hedge disappears and a working note starts looking like an approved product.
  • constraint Teaching service members that LLMs are uncertain puts the burden on whoever reads the document, and a reader cannot grade a claim whose origin is not printed on the page in front of them.
  • exposure Vendors selling models into the Pentagon are exposed to how their output gets rendered and forwarded. That exposure lands in support threads.

Two calls to the chatbot went into that report. Only the first one was analysis [22]. The analyst asked the tool to synthesize open-source data with classified signals intelligence. It misidentified the ship's cargo manifest [12]. The analyst then went back and asked it to format the erroneous finding into an official-looking summary, and that document circulated across command channels [13]. The second call added no fact and changed how the claim looked. A synthesis with a hedge in it looks like one analyst's working note. The same sentence in the house summary format looks like the command's position. The report said the ship was delivering components to a nuclear program in Iran, a country the US and Israel have been bombing since Feb. 28 [19]. By the time anyone established where the claim came from, military aircraft were in the air, and the operation was aborted at the last minute [11]. Officials found out the assessment was AI-generated only after it had nearly led them to apprehend the vessel [5]. TechCrunch reported that the concern among military officials and outside experts is that the errors these systems produce can travel up the chain of command before being questioned [14]. The Pentagon's own case for the tools is speed. It has described AI as delivering a significant advantage in speeding up its kill chain so commanders can respond in the right time [15]. Jake Steckler, a research scholar at GovAI and a veteran US Army officer, told TechCrunch that "It's important for service members to understand the uncertainty inherent to LLMs" [16], and that this is "especially critical for any decisions that could lead to use of force, like targeting, intelligence analysis, or operational planning. There are life and death consequences for those decisions" [17]. Steckler does not treat the episode as a reason to back away from the software. "But prioritizing adoption speed over all else will likely lead to incidents that only make service members lose trust in these systems, which ultimately is only going to slow adoption," he said [18]. CNN was not able to establish what the ship was actually carrying, or which AI tool the analyst used, commercial or built for the government [6]. The naming matters: only inside the department can anyone check whether the tool labeled its own output. The buying has not slowed in the meantime. Defense Secretary Pete Hegseth's memo directing the department to become "AI-first" is from January, and the aircraft were airborne in the spring [8][24]. A deal to let the military use Grok was reported in February, and NVIDIA, Microsoft and Amazon entered a DoD partnership in May [10]. Anthropic refused to let the military use its models to develop autonomous weapons and was briefly banned by the government [9]. There is a costlier version of the same workflow already on the record. Bloomberg reported that when US forces launched a missile at an Iranian school on the first day of the war, intelligence analysts had relied too heavily on Palantir's Maven Smart System, a tool that fuses dozens of data inputs, and the strike on the Shajarah Tayyebeh Elementary School killed more than 150 people, including 123 children [20]. UN experts said there were "reasonable grounds" to say the US had committed war crimes in its strikes against the school, according to the Associated Press [21]. The test that transfers to ordinary software is narrow. It comes down to the artifact your model produces that a person forwards without editing, and whether the provenance survives the forward. Provenance survives when the line sits inside the document, because the banner in the chat window stays in the chat window.

What to watch

  • Whether the Department of Defense responds to CNN's reporting or identifies the chatbot the analyst used.
  • The UN experts' findings on the Shajarah Tayyebeh Elementary School strike, due to be presented on Monday to the UN Human Rights Council.
  • Any DoD rule requiring generated intelligence products to be labeled as model output on the document itself, not just in the chat session.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories