A growth agency counted the articles. OpenAI has described the flaw in its own words. South African courts have been handed citations to cases that could not be found. The question is whether anyone checked.
In the spring of 2026, during the US-led war on Iran, an analyst at US Special Operations Command Pacific in Hawaii queried an AI chatbot about a Chinese cargo vessel operating in the Middle East. The chatbot examined the ship’s manifest, combined open-source intelligence with classified signals intelligence, and produced a finding: the vessel was carrying components of a nuclear weapons programme, bound for Iran.
The analyst used the same chatbot to format the finding into a standard intelligence report and circulated it through command channels.
The report was entirely false.
The ship was on a legitimate commercial voyage. It carried no nuclear material. The chatbot had hallucinated a conclusion from the data it was given, and the hallucination had travelled from a query on a screen in Hawaii into the operational pipeline of the US military.
Armed boarding teams prepared to intercept the vessel. Military aircraft were already in the air. Four sources familiar with the incident told CNN the operation was stopped only when officials examined the analyst’s report more closely and determined that AI had been used to generate its contents. They figured out that the chatbot had misidentified the cargo, and halted the operation.
One source described the report as “entirely false.” Another said it “almost started a war.”
CNN broke the story on 18 September 2026. Within days it was reported by Engadget, Gizmodo, the Jerusalem Post, the Taipei Times, the Times of Israel, the Maritime Executive, the Daily Beast, IBTimes, and TechTimes, among others. As of the date of this folio, neither the Pentagon nor Special Operations Command Pacific has commented on the specifics. CNN was unable to determine what AI tool the analyst used — whether commercial or government-made — or what the ship was actually carrying.
What the analyst did #
The sequence matters, because every step involved a person.
From query to aircraft: how a hallucination travelled
Reconstructed from CNN’s reporting and secondary sources · spring 2026
- Analyst queries AI chatbotAsks about a Chinese vessel’s cargo manifest, feeding it open-source intelligence and classified signals intelligence.
- Chatbot returns false assessmentConcludes the vessel is carrying components of a nuclear weapons programme, bound for Iran. The conclusion does not follow from the inputs.The hallucination is generated here.
- Analyst uses chatbot to format the findingThe same chatbot structures the false assessment into a standard intelligence report.
- Report circulated through command channelsThe formatted report enters the US military’s operational pipeline.
- Military prepares to interceptArmed boarding teams assemble. Military aircraft are launched. The operation to board the Chinese vessel is under way.
- Officials catch the errorSeparate officials examine the report, determine AI generated its contents, identify the cargo misidentification, and halt the operation.The catch came “within minutes” of the boarding preparation.
As TechTimes reported, quoting sources familiar with the incident: “A human was present at every stage. But the role of that human was to query and transmit, not to independently verify.”
This is the pattern Folio 05 described in a different setting. In South African courts, the Charlotin database records cases where litigants used AI to draft filings and submitted the results without checking them. The AI invented case law. The litigants transmitted it. The courts caught it.
The mechanism is the same. A person asks the machine. The machine answers confidently. The person passes the answer along. Based on what has been reported, nobody between the query and the submission checked whether the answer was true.
The difference is what was at stake. A false court filing wastes a judge’s time and damages a litigant’s case. A false intelligence report nearly put armed US forces on a Chinese vessel during a war between nuclear-armed powers.
Why the AI got it wrong #
Based on the publicly available information, no source has identified the chatbot involved. CNN could not establish whether it was a commercial product, something built for the military, or the Pentagon’s own GenAI.mil platform. That gap is itself significant: it suggests the military may have deployed AI tools across commands without a centralised record of which tools were in use or how they handled classified material.
What the sources do describe is how the chatbot processed the data. It fused open-source intelligence — publicly available information — with secret signals intelligence. From that fusion it generated a conclusion about the ship’s cargo that did not follow from the inputs. The conclusion was, in the language OpenAI uses to describe the same problem in its own models, a hallucination: a plausible but false statement generated by a language model.
OpenAI’s September 2025 research on why models hallucinate offers an explanation. Models “hallucinate because standard training and evaluation procedures reward guessing over acknowledging uncertainty.” A model graded only on right answers learns that a confident guess scores better than “I don’t know.”
That is what appears to have happened here. The chatbot was asked a question. Instead of flagging uncertainty — instead of saying it could not determine what the cargo was — it produced a specific, confident, and false answer. Nuclear weapons components. It is the kind of answer a model generates when it has learned that specificity is rewarded and uncertainty is not.
How close it came #
The operation was further along than a plan on paper.
Military aircraft were already airborne. Armed service members were preparing to board the vessel. Two sources told CNN about the boarding preparation. The operation was halted when officials — separate from the analyst who generated the report — examined the intelligence more closely and identified that AI had produced the assessment.
One detail reported by TechTimes: the catch came “within minutes” of the boarding preparation. The margin was not days. It was not hours.
The implications of a boarding were not abstract. The vessel was Chinese. Boarding a Chinese ship during a war — on intelligence that turned out to be fabricated by an AI — could have triggered an armed confrontation between the United States and China. That is what the source meant by “almost started a war.”
The incident was reported six days before Chinese President Xi Jinping’s state visit to Washington on 24 September 2026. AI governance was on the summit agenda — the two sides agreed to establish an AI safety incident hotline. No public Chinese reaction to the ship incident has been recorded in any source we found.
The policy that made it possible #
The incident did not happen in a vacuum. It happened inside a policy environment that was pushing AI adoption fast and had not, based on what sources described, built matching safeguards.
In January 2026, Defence Secretary Pete Hegseth launched the Pentagon’s Artificial Intelligence Acceleration Strategy. The strategy directed the Department of Defence to become an “AI-first warfighting force across all domains.” It called on the military to “unleash experimentation” and “eliminate legacy bureaucratic blockers” to AI integration.
What the strategy did not do, according to the sources CNN spoke to, was establish unified standards for how AI tools should be used across commands. Different branches and agencies were using different AI tools under different standing orders. According to those sources, there was no single set of rules for when AI-generated output required independent human verification before it could be acted on.
Sources told CNN that comparable errors had happened before. The Chinese ship incident was, in their words, “part of a trend in the US military, not an isolated incident.”
Could a tool have caught it? #
Based on what has been reported, the officials who halted the operation appear to have done something simple: they looked at the report and asked two questions — was this generated by AI, and does the claim hold up? Those two questions map onto categories of tools that exist today. The question is whether any of them could have been in the room.
Four layers of verification — and where the tools sit
Tools that exist today, mapped against what the military needed · September 2026
Layer 1: Was this AI-generated?
Tools like Turnitin, GPTZero, Copyleaks, and Originality.ai are built to answer this question. They examine text and estimate how likely it is to have been written by a machine. Turnitin checks whether a document was copied or AI-generated. GPTZero detects AI-written text and includes a hallucination check that verifies whether cited sources exist. Copyleaks identifies AI-generated and plagiarised text across languages. Originality.ai combines AI detection with a fact-checking feature it describes as a “fact checking aid.”
This is, we believe, roughly what the officials did by hand. They examined the report, determined that AI had generated its contents, and that recognition prompted them to check further. An AI-detection tool embedded in the intelligence pipeline could have flagged the report before it reached the operations room. But detection alone does not tell you whether the content is true. It tells you to look harder.
Layer 2: Does the output follow from the inputs?
A different class of tools checks whether an AI’s conclusions are actually supported by the documents it was given — a property researchers call groundedness or faithfulness. Patronus AI’s Lynx is an open-source model fine-tuned specifically for this task; it outperforms GPT-4 as a hallucination judge on published benchmarks. Galileo’s Luna does the same work at lower cost and higher speed than using a frontier model as an evaluator. Braintrust, Arize Phoenix, DeepEval, and TruLens offer similar scoring in developer platforms.
These tools are directly relevant to what went wrong. The chatbot was given a cargo manifest and signals intelligence. Its conclusion — nuclear weapons components — did not follow from those inputs. A groundedness checker sitting between the chatbot’s output and the formatted report could have flagged the gap: the inputs said one thing, the conclusion said another.
The limitation is practical, not technical. Every one of these tools is built for the software development pipeline — testing AI applications before deployment, monitoring production outputs, scoring retrieval-augmented generation systems. None is designed to operate inside a military command’s reporting chain, on classified networks, under the time pressures of an active operation. The concept is sound. The deployment environment does not exist.
Layer 3: Can the claim be confirmed?
A third category goes further than checking whether the AI was faithful to its inputs. It asks whether the claims in a document can be verified against independent sources.
Factiverse, used by Nordic government customers and independently tested at a NATO-certified security centre, checks claims against user-configured trusted sources and categorises them as supported, disputed, or mixed. ClaimBuster, developed at the University of Texas, identifies which claims in a text are check-worthy — a triage tool that tells you where to focus, not whether a claim is true. Originality.ai’s fact-checker sits in this space as well, though the platform describes it as an aid rather than a final authority.
Ukweli operates here. It reads a document before publication and checks whether the claims, quotes, and sources in it can be confirmed. It does not ask whether the document was AI-generated — that is the question the other tools ask. It asks whether what the document says is true, precise, and defensible.
In the military case, a claim-verification tool pointed at the analyst’s report would have found that no independent source confirmed the nuclear-cargo assessment. That is a useful signal — an unconfirmable claim in an intelligence report is a claim that needs a second look. But every tool in this category, Ukweli included, checks against sources it can reach. A classified cargo manifest is not in any public database. The ground truth was locked inside the same classified environment where the hallucination was generated.
Layer 4: Where did each sentence come from?
The deepest gap is provenance. When the analyst reportedly fed open-source intelligence and classified signals intelligence into the same chatbot prompt, the two became one undifferentiated stream of text. The output carried no record of which conclusion came from which source. As the AI research firm Pebblous noted in an analysis of this incident, “the moment open-source material and classified signals intelligence go into the same input box, the two become one stretch of tokens” and sentences carry no record of origin.
No deployed product tracks provenance at this level for intelligence workflows. The concept exists in research — source-attributed generation, confidence-scored sentences, grade-marked outputs that distinguish classified from open-source material. But no product integrates it into a classified pipeline. The eight AI firms cleared for the Pentagon’s classified networks in May 2026 — Amazon Web Services, Google, Microsoft, NVIDIA, OpenAI, Reflection, Oracle, and SpaceX — are providing general-purpose AI capability. None has announced a verification layer.
What this tells us
The tools that could have detected the hallucination in principle — groundedness checkers like Lynx and Luna — exist and are technically mature. The tools that could have flagged the claim as unconfirmable — verification tools like Factiverse, ClaimBuster, and Ukweli — exist and are in production use. The tools that could have traced the false conclusion back to its absence in the source material — provenance systems — do not exist as products.
Based on what has been reported, what caught the error was none of these. It was people doing what the analyst appears not to have done: reading the report critically and asking where the claims came from. The tools described above automate parts of that question. But in September 2026, based on all available reporting, none of them was in the room.
The reaction #
Three Democratic senators — Mark Warner, Jack Reed, and Chris Coons — wrote to Hegseth after CNN’s report, requesting an investigation. Their letter cited “growing concern about the extent to which agencies under your oversight have prioritised acceleration of AI capability adoption.”
The Pentagon’s public response was limited. It acknowledged using “rigorous verification of all intelligence inputs” but offered no specific reforms. Neither the Pentagon nor Special Operations Command Pacific responded to CNN’s requests for comment on the specifics of the incident.
No disciplinary action against the analyst has been reported. No policy change has been announced. The AI tool involved has not been identified or, as far as public reporting shows, withdrawn.
What this means #
This folio is about one incident. But the incident sits inside a pattern that the Integrity Watch series has been documenting.
Folio 05 showed that about half of new English-language articles online are now mainly AI-written, and that the tools producing them carry a known rate of false output — hallucinations that their own builders describe as a fundamental challenge. The EBU found that almost half of AI answers about the news had significant issues. The Lancet audit found fabricated references climbing steeply in the scientific record. South African courts are seeing invented case law in filings.
The Chinese ship incident is the same problem in a setting where the consequences are immediate and physical. A chatbot hallucinated. A person, according to sources, transmitted the hallucination without verifying it. The hallucination nearly triggered a military operation against a foreign vessel during a war.
The lesson is not that AI should never be used in intelligence. The lesson is the one that runs through every entry in this series: the output of these tools must be verified before it is acted on. In publishing, that means checking claims before they go out. In courts, it means checking citations before they are filed. In military intelligence, it means checking an assessment before it reaches the operations room.
The analyst in Hawaii, according to sources, did not check. The system around the analyst does not appear to have required checking. The policy environment was oriented toward speed, not verification. The people who caught the error appear to have done so by doing what the analyst had not: they looked at the report and asked where it came from.
That question — where did this come from, and is it true — is the one this series keeps returning to. It is the question Ukweli was built to help answer.
How we checked #
This folio draws on reporting by CNN and secondary coverage by eleven other outlets. Here is how we handled each source.
CNN’s exclusive report. The original story was published on 18 September 2026 by Zachary Cohen and Katie Bo Lillis. It is based on four sources familiar with the incident. CNN was unable to determine the specific AI tool used or what the ship was actually carrying. We were unable to access the CNN article directly (blocked by robots.txt) and relied on its details as reported by the secondary sources, all of which attributed the original reporting to CNN.
Secondary sources. We read reporting from Engadget, Gizmodo, the Jerusalem Post, the Taipei Times, the Times of Israel, the Defence Post, the Maritime Executive, the Daily Beast, IBTimes, TechTimes, and the World Socialist Web Site. We cross-referenced claims across sources and used only details that appeared consistently or were attributed to named sources. Where sources disagreed or added detail not in others, we noted the source.
Quotes. The direct quotes — “entirely false,” “almost started a war,” “a human was present at every stage” — appear across multiple outlets and are attributed to sources who spoke to CNN. We reproduce them as reported.
The Hegseth AI strategy. Details of the January 2026 Artificial Intelligence Acceleration Strategy come from the Daily Beast’s reporting and cross-referenced against the WSWS article. We describe the strategy’s directives as reported.
The Congressional letter. The letter from Senators Warner, Reed, and Coons is reported by the WSWS and referenced in the military.com article (which we could not access). We describe its contents as reported.
The verification tools survey. The section “Could a tool have caught it?” draws on published product descriptions, benchmarks, and technical documentation from each named tool. Patronus AI’s Lynx benchmarks are from its published research paper and GitHub repository. Galileo Luna’s performance claims are from VentureBeat’s reporting and the paper published at COLING 2025. Factiverse’s capabilities and NATO certification are from the company’s website. The Pebblous provenance analysis is from the firm’s published commentary on this incident. The eight Pentagon-cleared AI firms are from Defense One’s May 2026 reporting. Tool descriptions for Turnitin, GPTZero, Copyleaks, Originality.ai, ClaimBuster, Braintrust, Arize Phoenix, DeepEval, and TruLens are from each product’s public documentation as of September 2026.
What we do not know. We do not know which AI tool was used. We do not know what the ship was actually carrying. We do not know whether any disciplinary or policy action followed. We do not know China’s official response. We do not know whether any verification tool was available to the analyst or the command. These gaps are stated in the folio.
Sources #
- Cohen, Z. and Lillis, K.B., “Exclusive: US military had close call after using AI for false intelligence report, sources say”, CNN, 18 September 2026.
- Campbell, I.C., “AI almost led the US military to start a war with China, report says”, Engadget, 19 September 2026.
- “‘Almost Started a War’: US Military Nearly Boarded a Chinese Ship Based on Bad Intel From AI”, Gizmodo, 19 September 2026.
- “US nearly launches military operation over AI-generated information”, Jerusalem Post, 19 September 2026.
- “Faulty intel almost started a war after AI misread Chinese vessel cargo: CNN”, Taipei Times, 20 September 2026.
- “US nearly raided Chinese ship after AI falsely flagged ‘nuclear’ cargo for Iran”, Times of Israel, 19 September 2026.
- “AI Error Nearly Sent US Forces to Intercept Chinese Ship”, The Defense Post, 21 September 2026.
- “Report: AI Mistake Nearly Led U.S. Military to Board a Chinese Ship”, Maritime Executive, 21 September 2026.
- “Pentagon Pete Hegseth’s AI Strategy Almost Started War With China”, The Daily Beast, 20 September 2026.
- “‘Almost Started a War’: False AI-Linked Intel Triggered US Military Moves Towards Chinese Ship”, IBTimes, 20 September 2026.
- “US military AI hallucination ‘almost started a war’ with China”, World Socialist Web Site, 22 September 2026.
- “US Military Almost Boarded Chinese Ship Over AI-Hallucinated Nuclear Claim”, TechTimes, 20 September 2026.
- OpenAI, “Why Models Hallucinate”, published 5 September 2025, at openai.com.
- Folio 05: “Half of new articles online are now mainly AI-written. The tools that write them hallucinate”, Ukweli, 2026, at ukweli.io/blog/folio-05.
- Patronus AI, “Lynx: State-of-the-Art Open Source Hallucination Detection Model”, at patronus.ai/blog/lynx.
- Galileo AI, “Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost”, COLING 2025.
- Factiverse, “AI and Media Produce Claims. We Check Them”, at factiverse.ai.
- Pebblous AI, “Military AI Hallucination and the Provenance Gap in Merged Reports”, at blog.pebblous.ai.
- “8 AI firms cleared to provide tools for classified Pentagon networks”, Defense One, May 2026.
The standing block #
Read the register and download the data →
Integrity Watch is a verified register, neutral by design. Every entry is confirmed against a reputable source, and if it cannot be confirmed, it is not listed. Ukweli is a content integrity platform built by The Orkestra.
Get the next folio
One email a month from Integrity Watch: the new folio, and what changed in the register since the last one. Nothing else, and you can leave at any time.
The register itself stays free and public, and the full dataset is yours to download whether you subscribe or not.
Know the truth before you publish.
Ukweli checks documents for fabricated and unverifiable citations, sources, quotes and claims, before they leave your hands. Integrify a document · Get the next folio