On Tuesday a generative artificial‑intelligence system mis‑identified a non‑existent threat, nearly prompting a US military strike before senior officers halted the order.
How the error unfolded
According to a TechCrunch report, the incident occurred at a Pentagon‑level command hub where a large‑language model, integrated into an intelligence‑analysis workflow, flagged a hostile activity off the coast of Yemen. The model’s output suggested an imminent attack, prompting the operations team to draft a launch order.
Within minutes a junior analyst noticed that the AI‑generated brief contained inconsistent geography and dates that did not match any known intelligence feed. The analyst raised the discrepancy, and a senior officer suspended the plan pending verification.
“It’s important for service members to understand the uncertainty inherent to LLMs," a GovAI research scholar warned.
The halt prevented the deployment of aircraft and the release of live ordnance. No casualties occurred, and the operation was called off before any assets left the base.
Official reaction
The Department of Defence released a statement saying the incident “highlights the critical role of human judgement in the loop when employing advanced AI tools.” It confirmed that a review of the AI system’s training data and decision‑making pipeline is under way.
GovAI, the government‑funded research programme that developed the language model, said its scholars will publish a set of safety guidelines aimed at quantifying model uncertainty and enforcing mandatory human validation for any action‑able recommendation.
Wider implications for defence AI
Generative AI has been fast‑tracked into US defence projects since 2022, with promises of faster threat analysis and reduced workload for analysts. However, earlier incidents — such as a 2024 false alarm where an AI‑driven radar system mistakenly identified a flock of birds as hostile drones — have already raised doubts about reliability.
Congressional committees are now demanding tighter oversight, and the Pentagon’s Joint Artificial Intelligence Center has pledged to tighten testing standards before any model is deployed in a live‑fire context.
Experts say the Tuesday episode underscores a broader risk: large‑language models can “hallucinate” details that appear plausible, especially when fed sparse or noisy data. Without robust uncertainty quantification, such hallucinations can translate into real‑world operational hazards.
The next step, officials say, is a multi‑agency audit of all AI‑assisted decision tools, coupled with mandatory training for service members on the limits of machine‑generated intelligence.
Discussion (0)
Sign in to join the discussion.