AI systems causing real harm before oversight can catch up
A hallucination in a military AI system nearly triggered a US attack on Chinese nuclear infrastructure. This week's developments, taken together, show a widening gap between what AI systems can do and what the humans overseeing them can actually catch.
Neo NeumannAI Practice LeadSeptember 18, 2026This week's news shares a common thread: AI systems producing consequential outputs, often at speed and scale, while the humans nominally in charge fall further behind. The incidents are not isolated glitches. They form a pattern worth taking seriously now, before your organisation becomes the next case study.
An AI hallucinationAI hallucinationA hallucination is when an AI model generates output that is fluent and confident but factually wrong, fabricated, or unsupported by its source data.View full definition → almost triggered a military strike
Ars Technica reported this week that a hallucinationhallucinationA hallucination is when an AI model generates output that is fluent and confident but factually wrong, fabricated, or unsupported by its source data.View full definition → in a military AI system nearly led to a US attack on Chinese nuclear infrastructure. The system generated false information about Chinese nuclear components, and the output was treated with enough credibility that an actual strike was considered before the error was caught.
The details of exactly how close this came to action remain unclear from the reporting, but the structure of the failure is clear enough: an AI system produced a confident, specific, and completely wrong assessment in a domain where being wrong carries catastrophic consequences.
What makes this worth stopping on is not just the severity. It is the mechanism. AI hallucinations in high-stakes work) are not about systems going rogue or malfunctioning in some dramatic way. They are about systems doing exactly what they were designed to do, which is generate plausible, fluent, confident-sounding text, and that output happening to be false. Military decision-making under time pressure is precisely the environment where the gap between "plausible" and "verified" is most lethal.
The same Ars Technica reporting notes that the military's overall use of AI appears to be accelerating, not pausing for reflection. If anything, the incident is being absorbed into a trajectory of expansion rather than promptingpromptingPrompt engineering is the practice of designing and refining text inputs to guide large language models toward accurate, relevant, and reliable outputs.View full definition → a structural reassessment.
For professionals not working in defence: the principle transfers directly. If your organisation uses AI outputs to inform decisions about credit, hiring, medical protocols, contract terms, or supply chain risk, the question is not whether hallucinations will occur. They will. The question is whether your process would catch them before they matter.
Do not wait for governance frameworks to mandate verification. Build it into your workflow now.
Researchers used Claude to hack OpenAI
Security researchers used Anthropic's Claude to exploit vulnerabilities in OpenAI's systems, gaining access to an employee account and an internal code repository, before reporting the flaws. The story was covered by Ars Technica, TechCrunch, and The Decoder, making it one of the more thoroughly reported incidents of the week.
The specifics matter here. This was not a theoretical red-team exercise. According to TechCrunch, the researchers took over employee accounts and reached sensitive GitHub data within 72 hours. The attack vector was an AI model, pointed at another AI company's infrastructure.
This raises a question that most security teams have not yet formally answered: if an adversary can use a publicly available LLMLLMA Large Language Model is an AI system trained on vast text data to predict and generate language, enabling tasks like writing, summarizing, and answering questions.View full definition → to probe and exploit your systems, does your current security posture account for that? AI-assisted attack is not a future risk category.
The irony of Anthropic's Claude being used to breach OpenAI will attract attention, but the actually transferable point is narrower. AI tools lower the skill floor for conducting sophisticated attacks. The timeline shrinks. The expertise required decreases. Your exposure window gets shorter, not longer.
Watch whether either company discloses what was accessed, and whether this accelerates security standards for AI system integrations more broadly.
The "AI overseeing AI" answer to agent oversight
TechCrunch reported this week that as companies deploy AI agentsAI agentsAgentic AI refers to AI systems that pursue goals autonomously by planning, taking actions through tools, and adapting based on results, with minimal step-by-step human direction.View full definition → on longer and more complex tasks, a new oversight problem has emerged: agents can act faster and at greater volume than humans can realistically review. The proposed solution, from several companies now working on this, is to use additional AI to monitor the agents.
The logic is not absurd. Human reviewers cannot read every output a system generates at production speed. But the reasoning deserves scrutiny. Using AI guardrails and human-in-the-loop design) has always been about preserving meaningful human judgment at the points where it actually matters, not replacing human attention with automated attention and calling the problem solved.
If an AI agent makes errors, and a monitoring AI fails to catch them, the error chain now has two AI-generated links before a human sees anything. That is not obviously safer. It may feel more manageable operationally, but operationally manageable and genuinely safe are different things.
For anyone deploying or evaluating agent systems in their organisation: ask specifically what the monitoring layer can and cannot detect. If the answer involves another model's judgment, understand the failure modes of that second model too.
The signal underneath: governance is not coming fast enough to matter
The hallucination near-miss, the AI-assisted security breach, the agent oversight gap, these are all, at their core, governance failures. Not failures of the technology to perform, but failures of the institutional structures around the technology to keep pace with what the technology is doing.
The EU AI Act exists, and its risk classifications are reasonable in principle. But the nuclear hallucination story illustrates the core problem: by the time a governance framework is written, debated, adopted, and enforced, systems with the capacity to generate that kind of failure are already deployed in critical environments. The gap between what regulators can legislate and what engineers can ship has never been wider.
The development most people are underrating this week is not any single incident. It is the aggregate signal that self-regulation by deploying organisations is currently the only functioning check that exists at speed. External oversight is years behind. Internal oversight is outpaced by deployment velocity. The question every manager should be asking their AI-adopting teams is: who in this organisation is responsible for catching the wrong answer before it becomes a decision?
If the honest answer is "the model's confidence level" or "we assumed it was checked," that is the gap to close first.
The military hallucination story is the sharpest version of a problem that exists in far more ordinary settings. Build verification into your process as a standing requirement, not a one-time audit. The organisations that treat this seriously now will have the institutional muscle for it when the stakes in their own domain turn out to be higher than they expected.
Go deeper
The lessons that take this article further, free to read.
- 1Hallucinations and verification in high-stakes workResponsible & trustworthy AI
- 2Hallucinations: why confident answers can be wrongAI & LLM foundations
- 3Governance and the EU AI Act: the toplineResponsible & trustworthy AI
- 4Guardrails, permissions, and human-in-the-loopAI agents: design, build & operate
- 5Verifying outputs: trust but checkAI in daily work
Sources
- AI hallucination of Chinese nuclear components almost led to US military attack
- A new kind of AI model from a ChatGPT inventor is thrilling developers
- Security researchers used Anthropic's Claude to hack OpenAI's internal systems in under 72 hours
- AI training built on fair use looks shaky when the companies' own people call it "astonishing theft"
- Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think
- Researchers used Anthropic’s Claude to hack into OpenAI
- New experts join Google’s AI & Economy team
- Researchers used Claude to hack OpenAI
- Co-creating the future of fashion with Google
- 5 Prompt Optimization Strategies That Actually Improve LLM Output
- PrismML hopes its tiny LLM will change how we all use AI
- The fix for rogue AI agents could be more AI
- Is the AI safety debate about safety or control?
- Making global data easier to explore
Finished reading?
Validate your read to earn XP and feed your radar.