At around 4am on 28th March 1979, a chain of failures began at Unit 2 of the Three Mile Island Nuclear Generating Station in Pennsylvania. Within minutes, more than a hundred alarms were sounding in the control room. The instruments incorrectly told the operators that a relief valve was closed. Confronted with what one of them later called “an avalanche of alarms” they struggled to distinguish signal from noise. They concluded that the system had too much water when in fact it was losing coolant and shut down the emergency cooling system that might have saved the reactor. The core suffered a partial meltdown. The cleanup took fourteen years and cost over a billion dollars. Craig Faust, an operator on the shift, later told the President’s Commission: “I would have liked to have thrown away the alarm panel. It wasn’t giving us any useful information.” The Commission concluded that the operators were overwhelmed by the sheer volume of information they were receiving. But it also revealed that some instruments were giving misleading information that made it hard to build a true picture of the state of the reactor. For example, the light on the control panel indicated that the command to close the relief valve had been sent but did not indicate whether the valve was actually shut.
In the age of AI, executives are starting to experience challenges analogous to those faced by the operators at Three Mile Island. The volume and cadence of information received is overwhelming. And that information, whilst appearing polished and convincing, is of variable quality and may provide unreliable signals about the true state of the business. And these two factors interact. The result is a plausibility crisis: more material reaches decision-makers looking polished, reasoned and complete, even when the underlying logic or evidence is weak.
Overwhelmed by Information
Today’s AI tools are increasing the volume and cadence of information across every channel into an executive. The time needed to produce a polished and well researched strategy, plan or business case has fallen from weeks to hours. And whilst AI tools can be used to improve brevity, the opposite is often true. Junior staff are often more fluent than their managers and use AI liberally to increase their volume and pace of output. They are ‘tooled up’ - 78% of employees admit to using unapproved AI tools. Customers and counterparties produce AI-generated requests, proposals and responses. Boards, simultaneously excited and anxious about AI, are adding to the pile.
The use of AI tools has also upped the ante on expected turnaround times. Historically it might have been reasonable to allow colleagues or suppliers days or weeks to turnround a proposal after delivering feedback. In today’s context that looks generous. As AI agents enter the workplace, we can expect turnaround times to further compress. An agent can process incoming email and send a response in seconds without human input, for example.
Executives have always had to deal with information overload, but this is overload of a different type. Renowned software engineer and blogger Steve Yegge notes that AI is automating the easy work and leaving only the hard decisions, at a pace, he says, anyone can hold for “a few hours once or occasionally twice a day, even with practice”. He calls it “the AI vampire”.
The trajectory is steep. Microsoft’s 2025 Work Trend Index, drawing on Microsoft 365 telemetry, found the average employee now receives 117 emails and 153 Teams messages a day with interruptions arriving every two minutes and 40% of workers opening email before 6 a.m. The volume problem shows no sign of plateauing.
Unreliable Signals
The second challenge for executives in the AI era is finding the signal in the noise. After decades of dealing with patchy management information, the executive layer increasingly has access to real-time, granular and accurate management information. The plausibility crisis begins at one level of abstraction up from the dashboard. AI doesn’t corrupt the underlying data, it corrupts the interpretation of the data. Generative AI often writes the analysis, summary or recommendation. Because the underlying numbers are correct, the output looks credible - perhaps more credible than the human output it replaced. So, the reader has no immediate visual cue that distinguishes solid analysis from hallucination. There is nothing visibly wrong to react to. As with a powerful orator, style can mask a lack of substance. It’s harder to find the deep insights and harder to spot the logical flaws. AI tools are improving, but they still make plenty of mistakes. The European Broadcasting Union and BBC’s October 2025 study had journalists from 22 public broadcasters evaluate 3,000 responses from LLMs, finding that 45% contained significant errors and 20% were judged unreliable. Vectara’s November 2025 refresh of its hallucination benchmark, using longer enterprise-realistic documents, found the latest reasoning models (GPT-5, Claude Sonnet 4.5, Gemini 3 Pro) all hallucinated more than 10% of the time.
AI hallucinations can have meaningful consequences. In February 2024, Air Canada was held liable in the British Columbia Civil Resolution Tribunal after its chatbot invented a bereavement-fare refund policy that did not exist; the airline was ordered to honour the chatbot’s fabrication. In October 2025, Deloitte Australia refunded part of an A$440,000 government report after academics identified fabricated citations and a quote from a Federal Court judgment that the judge had never written.
Compounding Workslop
These two problems do not add, but rather multiply. As executives become overloaded dealing with a higher volume and cadence of input, the natural response is to reach for AI tools to keep up. The output goes out unverified, because taking the time to verify it would defeat the point. Recipients in turn then verify with their own AI because they too are overloaded. Both ends may use similar models, tools or training corpora so both miss the same errors. Two negotiating teams prepping with the same AI model are more likely to miss the same flaws in the contract.
The immediate impact is more local. AI content generation gets easier and cheaper every quarter as the LLMflation curve continues to compound. But verification – finding and correcting the errors introduced by AI – gets more expensive. The cognitive cost to check each artefact remains constant. Users of AI may report a productivity gain. But few can measure or perhaps even see the underlying errors and risks.
Five Fixes
Deployed well, AI is a powerful tool, and one that no organisation can afford to overlook. What steps can firms take to simultaneously boost AI uptake and return on investment, whilst mitigating the challenges of information overload and weak signal? Paradoxically, the solution involves increased use of AI. But rather than asking individual executives to become better prompt engineers, AI needs to be integrated into newly configured end-to-end workflows. In this way, volume is cut, signal and noise are more cleanly separated, and judgement is deployed where it is really needed. Five specific changes to the operating model are needed:
1. Establish an integrated ‘centaur’ (human and machine) workflow. The point is not to let one model check another. It is to separate generation, checking, routing and escalation so that scarce human judgement is applied at the right point. This requires an integrated workflow that captures all human and machine actions, criteria and routing. Not all decisions can or should be verified by AI. But low-risk content can be triaged this way so that only the more important decisions make their way to the executive inbox. Goldman Sachs rolled out GS AI Assistant firmwide in June 2025, with adversarial review built into the workflow. The US Treasury’s February 2026 Financial Services AI Risk Management Framework points toward risk-tiered governance, validation, monitoring and accountability for AI/model use.
2. Increase human and machine diversity. Two LLMs trained on overlapping corpora produce correlated errors. They tend to be wrong about the same things in the same ways. Two ICML 2025 papers (Kim et al.’s evaluation of more than 350 LLMs, and Goel et al.’s “Great Models Think Alike”) found that frontier models produce correlated errors. On one leaderboard, models agreed 60% of the time when both were wrong. The benefits of human diversity are well documented. Machine diversity can be increased by using different model families, providers, retrieval sources, evaluation methods, prompts and review workflows. In the highest-risk cases, firms may also need to establish separate teams and training data.
3. Deploy schemas. The human brain finds it easiest to process information that arrives in a familiar format (this is why most TV remote controls look alike). Schemas shift cognitive load from the reader to the writer as the reader expends less brainpower on parsing and more on evaluating. More importantly, schemas can explicitly surface implicit assumptions or weaknesses that otherwise would remain hidden. Most investors use a templated investment committee memo with a section on risks, as that forces a discussion on the downsides of the investment. In a similar way, AI (ideally a different model) can be made to surface assumptions, risks and logic gaps.
4. Separate one-way and two-way doors. Jeff Bezos’s Type 1 / Type 2 decision framework separates reversible and irreversible decisions. Reversible decisions can be taken through fast lanes, where AI can be deployed more liberally. Irreversible decisions need human deliberation and verification. Central banks offer a useful analogy: scheduled rate decisions go through structured deliberation, while routine market operations are executed through faster, pre-authorised processes. Importantly, criteria for which decisions go through which lane need to be well defined and understood.
5. Close the loop. AI models are evolving all the time. Models that make mistakes today will typically make fewer mistakes tomorrow. But that is not always true. OpenAI’s o3 hallucinated on 33% of PersonQA questions, double the rate of its predecessor o1 (16%). And newer models can be wrong in different ways. All AI-generated artefacts should be stamped with provenance including the prompt, model version and accountable human. Then, audit all AI models in the inventory periodically to understand usage, accuracy, return on investment etc. Based on the audit, models can be upgraded or retired. And the centaur workflow can be reconfigured to optimise the volume and increase the signal to noise ratio.
Questions for the Board
The board’s role is to oversee the enterprise risk presented by the plausibility crisis. It needs to be able to size the firm’s exposure and to oversee the work being done to reduce it. Five diagnostic questions can help shape the discussion:
1. Where is AI-generated content creating risks for our organisation?
2. How are we balancing the upside from AI productivity with the downsides of executive overload and signal confusion?
3. How diverse are our human and machine verifiers, and how do they combine?
4. What’s our process for auditing our models, and how are those audits being used to improve our operations?
5. Is the board getting reliable information? How much of the board pack was AI-generated and what human verification has taken place?
The Last Line of Defence
At Three Mile Island, the operators wanted to throw away the alarm panel. The executive layer cannot. What it can do is redesign what reaches it, who checks it, and which decisions it is allowed to deliver to the board unverified. AI did not create the plausibility crisis. The operating model did.
Footnotes & Sources
• Three Mile Island, Unit 2, partial core meltdown, 28 March 1979. Report of the President’s Commission on the Accident at Three Mile Island (Kemeny Commission), October 1979; US Nuclear Regulatory Commission, Backgrounder on the Three Mile Island Accident. Craig Faust testimony to the President’s Commission. Cleanup completed 1993; total cost in excess of $1bn.
• WalkMe / IDC State of AI in the Workplace, Propeller Insights for WalkMe, August 2025 (n=1,000 US working adults). 78% of employees admit to using unapproved (“shadow”) AI tools.
• Steve Yegge, “The AI Vampire,” Medium, 11 February 2026.
• Microsoft Work Trend Index 2025, “The Frontier Firm Is Born” and “Breaking Down the Infinite Workday.” Edelman Data x Intelligence, n=31,000 across 31 markets, plus Microsoft 365 telemetry. 117 emails plus 153 Teams messages received per employee per day; interruptions every 2 minutes; 40% of workers checking email before 6 a.m.; meetings starting after 8 p.m. up 16% year on year.
• Gartner forecast on AI-generated outbound messaging: Gartner press release, “Gartner Expects 60% of Seller Work to Be Executed by Generative AI Technologies Within Five Years,” 21 September 2023. 30% of outbound marketing messages from large organisations to be synthetically generated by 2025, up from less than 5% in 2022.
• European Broadcasting Union / BBC, News Integrity in AI Assistants Study, October 2025. Professional journalists from 22 public service media organisations across 18 countries and 14 languages evaluated 3,000 responses from ChatGPT, Copilot, Gemini and Perplexity. 45% contained at least one significant error; 31% had serious sourcing problems; 20% were judged completely unreliable.
• Vectara Hallucination Leaderboard, late-2025 refresh, “Introducing the Next Generation of Vectara’s Hallucination Leaderboard,” 19 November 2025. New benchmark uses longer enterprise-realistic documents (up to 32K tokens) spanning law, medicine, finance, technology and education. GPT-5, Claude Sonnet 4.5 and Gemini 3 Pro all exceeded 10% hallucination on grounded summarisation.
• Moffatt v. Air Canada, 2024 BCCRT 149, British Columbia Civil Resolution Tribunal, 14 February 2024. Tribunal found Air Canada liable for negligent misrepresentation after its chatbot incorrectly advised the claimant that bereavement-fare refunds could be claimed retroactively; airline ordered to honour the chatbot’s representation. Damages awarded: C$812.02.
• Deloitte Australia partial refund to Department of Employment and Workplace Relations, October 2025. A$440,000 (~US$290,000) report on the Targeted Compliance Framework contained fabricated academic citations and a fabricated quote from a Federal Court judgment; revised version disclosed use of Azure OpenAI GPT-4o tooling. Partial refund of approximately A$97,000. Sources: Australian Financial Review, Fortune, CFO Dive (October 2025).
• LLMflation: a16z, “Welcome to LLMflation: LLM Inference Cost is Going Down Fast,” and Stanford HAI, 2025 AI Index Report.
• Goldman Sachs GS AI Assistant: firmwide rollout 23 June 2025. Sources: Reuters, “Goldman Sachs launches AI assistant for its bankers, traders and asset managers,” June 2025; PYMNTS, “Inside Goldman Sachs’ Big Bet on AI at Scale,” 2025.
• US Treasury Financial Services AI Risk Management Framework, February 2026, and revised US banking model-risk guidance. Together, they point toward risk-tiered governance, approval and monitoring expectations for AI in financial services.
• Goel, Struber, Auzina et al., “Great Models Think Alike and this Undermines AI Oversight,” ICML 2025 spotlight; arXiv:2502.04313, February 2025. Finds that as frontier model capabilities increase, model errors become more correlated — raising risks of correlated failures in AI oversight settings.
• Kim, Garg, Peng & Garg, “Correlated Errors in Large Language Models,” ICML 2025; arXiv:2506.07962. Large-scale empirical evaluation of more than 350 LLMs. On one leaderboard dataset, models agree on the same wrong answer 60% of the time when both err.
• Bezos two-decision framework: Jeff Bezos, 2015 Letter to Shareholders, Amazon. Type 1 decisions are consequential and irreversible (“one-way doors”) requiring deliberation; Type 2 decisions are reversible (“two-way doors”) that can be made quickly.
• OpenAI o3 and o4-mini system card, April 2025. PersonQA hallucination rates: o1 (predecessor): 16%; o3-mini: 14.8%; o3: 33%; o4-mini: 48%. OpenAI: “more research is needed” to understand why hallucinations are getting worse as reasoning models scale up. Coverage: TechCrunch, “OpenAI’s new reasoning AI models hallucinate more,” 18 April 2025.

