In Case You Missed It
The Cost of AI Is Collapsing. Can Your Organisation Respond? Inference costs are falling tenfold a year, yet the bottleneck has simply moved from the model to the operating model.
Is AI Blunting Your Strategy? As AI automates the analysis, it quietly erodes the conditions that build judgement, and convergence on the same models breeds strategic monoculture.
Double (AI) Agents Autonomous agents act inside the business like recruited insiders, so governing them owes more to intelligence tradecraft than to IT.
The de Havilland Comet, the first jet airliner, was in its time the most advanced aircraft in the world. But on 10 January 1954 a Comet left Rome and broke apart at 27,000 feet. Three months later a second Comet disintegrated near Naples, and the fleet was grounded. Investigators working at Farnborough pressurised a fuselage in a water tank until, after some 3,000 simulated flights, the cabin tore open at a window corner. Metal fatigue had led to structural failure of the fuselage. Jet performance had outrun the assumptions built into the airframe.
The fix was to redesign the airframe to meet the operating demands of jet flight. Something similar is happening with AI. The models race ahead, but the organisations deploying them do not. Inference costs have fallen more than 280-fold between November 2022 and October 2024, yet around 95 per cent of corporate AI pilots produce no measurable profit, and the share of firms abandoning most AI initiatives before production has jumped from 17 to 42 per cent in a year. The better AI becomes, the harder it is to deploy.
Not Like the Last One
It is tempting to dismiss this as the usual lag seen with every general-purpose technology. Electrification took four decades to show up in the productivity figures. Early factories ran every machine off one central steam engine through a maze of overhead shafts and belts, so layout was dictated by the drive shaft rather than the work. When electricity came, most firms just swapped in one big motor to turn the same shafts, and the initial productivity gain was marginal. The uplift came only when they put an individual motor on each machine and rebuilt the factory around it, reordering the line around the flow of work and remaking the building, the skills and the management to match.
But that rebuild was only possible because the technology held still. Productivity lags close when a technology’s interfaces, to both people and other machines, and its operating assumptions, such as performance and error rate, are stable enough to rebuild the organisation around it. That stability is what let firms commit to a redesign and capture the returns.
Artificial intelligence weakens that precondition. The model is probabilistic, so it never settles into behaving the same way twice, and it keeps changing, so the target you would rebuild around keeps moving. The rebuild is never finished.
Compounding Challenges
More capable models can be less trustworthy. A more capable model is more fluent and more plausible, so its errors do not vanish, they hide, and the Verification Tax rises rather than falls. OpenAI’s newer reasoning models fabricated more, not less, on one benchmark about real people: invented claims climbed from 16 to 48 per cent. In 2026 EY Canada withdrew a cybersecurity report after most of its cited sources proved fabricated, attributed to outlets that never published them. New York City ran an official chatbot that gave unlawful guidance for two years, on tenant discrimination and on skimming staff tips, at a cost of around half a million dollars.
This is made worse by how people respond to a better model: they trust it more and check it less, just as its mistakes turn rarer and more costly. Lisanne Bainbridge named this “the irony of automation” in 1983: automate most of a task and you leave the human only the part you could not automate, let their skill decay, and still need them to step in the moment it fails. The erosion is measurable. Across nineteen experienced endoscopists, the rate at which they detected adenomas without AI fell from 28 to 22 per cent after their clinics adopted it.
And as the model is trusted more, the blast radius grows. A weak model drafts an email. A capable one is handed an agent with a level of autonomy. In July 2025 Replit’s coding agent reportedly deleted a live production database during an explicit code freeze.
More capable models require more organisational change to deploy. Material enterprise value is usually released only when you rebuild the organisation around the model. McKinsey finds that redesigning workflows is the single largest driver of profit from AI, yet only about a fifth of firms have redesigned any. Taco Bell, having put voice AI in more than 500 drive-throughs, said in 2025 it was rethinking where to use it and putting people back on the busiest lanes. The more the model can do, the more of the company it forces you to rebuild.
The Increasing Rate of Change
These four challenges compound as the model grows more capable. A fifth comes from the rate of change, and it resets the others before they resolve. Each upgrade invalidates the tests that cleared the last model and shifts the failures your people had learned to spot. In April 2025 OpenAI shipped an update to GPT-4o, found it had turned sycophantic, and pulled it within days. OpenAI later said it lacked specific deployment evaluations for the trait. Anthropic retires older Claude models on a regular cadence, each one pushing dependent workflows onto a successor that has to be re-validated. You cannot industrialise a process on ground that re-platforms every quarter.
The gap between capability and capture is not idle. It is accumulating Stranded Capability, capability you can already buy but cannot convert, and it compounds, because each upgrade multiplies the workflows, tests, controls and behaviours that must be relearned. Electrification paid off only once the technology was stable enough to redesign around. If the model never stays stable, the countdown keeps restarting.
Two objections could be made to this argument. First, better models should reduce error, so verification ought to get easier. Sometimes they do: OpenAI reported lower hallucination rates for GPT-5 than for earlier models. But verification does not disappear. The errors that remain can be more fluent, more context-specific and harder to spot at the edges. Second, the leaders are capturing value, so this is an execution issue, not a broad paradox. But the leaders are a minority, the roughly 6 per cent of high performers in McKinsey’s data who went furthest in rebuilding their workflows. They invested in the deployment machinery because the model alone did not deliver. They are the proof of the rule rather than the exception to it.
Building the Deployment Machine
None of this argues for waiting. It argues for spending the effort where the constraint is.
• Fund the deployment machinery, not only the model. Balance the budget between the model capability you rent and the in-house capability (skills, processes, etc.) that you build.
• Select and scope use cases purposefully. Select based on a deliberate mix of impact, reversibility, visibility and required organisational change. Begin where errors are cheap to undo, and defer the visible, irreversible cases until the machinery is proven.
• Build an abstraction layer. Put a gateway between your systems and the model so you can switch or multi-source without rebuilding everything. Test every upgrade against your own tasks rather than the vendor’s benchmark and run the upgraded model on a representative sample before broad deployment.
• Engineer the human backstop. Keep people in the loop where the consequences are high, rotate and sample the work so complacency cannot set in, and re-educate your reviewers on the new failure modes after every upgrade, because they were trained on the old ones.
• Measure the right things. Track three things: outcomes from AI deployment (e.g., cycle time, error rate, yield, process productivity, customer satisfaction); capability measures (e.g., share of use cases under monitoring, processes genuinely re-engineered); and model health (e.g., drift, error rates). Counts of seats and prompts measure activity but not value delivered or capability built.
• Rebuild the accountability framework. AI failures are often governance failures as much as technical failures. Name who owns selection, model changes, the human backstop, the outcomes and the risk.
Questions for the Board’s Back Pocket
• Are we allocating resources and funding optimally? Is most of our budget buying model capability we rent, when the true constraint is the conversion machinery we must build?
• Are we increasing or reducing strategic optionality? Could we switch models or vendors without re-platforming, or are we hard-wiring today’s frontier and sinking capital into something that will move under us?
• Do our metrics drive the right behaviours? Do we reward the business outcome and the building of the machinery, or do we reward usage, seats and prompts?
• Are we preserving and improving human judgement? As we trust the model more and check it less, can our people still catch what it gets wrong, and do we re-educate them each time it changes?
• Is our accountability framework fit for purpose? Is ownership senior and explicit, or has it defaulted to IT, and does our governance fit a substrate that changes every quarter rather than the deterministic projects it was built for?
The jet age scaled when aircraft makers rebuilt the whole aircraft around the operating demands of jet flight. The AI age will scale the same way: when enterprises build deployment machinery that accounts for what the models get wrong and stays resilient as they keep changing. The model was never the hard part.
Footnotes and Sources
• Cohen Inquiry, de Havilland Comet. Court of Inquiry report C.A.P. 127 (1955). BOAC Flight 781, 10 January 1954, all 35 killed; second Comet loss near Naples, 8 April 1954, 21 killed. RAE water-tank testing identified metal fatigue around window/opening structures. Comet 4 opened regular jet-powered transatlantic service in 1958; the Boeing 707 entered transatlantic service shortly afterwards and later won the larger market.
• Stanford HAI, AI Index 2025. The cost of reaching a fixed quality threshold fell more than 280-fold (about $20.00 to $0.07 per million tokens) between November 2022 and October 2024; SWE-bench scores rose from 4.4% to 71.7% in a year.
• MIT Media Lab, Project NANDA, “The GenAI Divide” (July 2025). Around 95% of enterprise GenAI pilots showed no measurable P&L impact. Caveat: 153-respondent senior-leader survey plus 52 interviews; the authors call it a “directionally accurate snapshot”; NANDA itself has an interest in agent infrastructure.
• EY Canada cyber report. Report withdrawn in 2026 after AI-detection firm GPTZero found 16 of 27 cited sources fabricated, misattributed or dead, with references falsely ascribed to Forbes, McKinsey, Gartner and others. EY confirmed the removal and a review. Source: Computing / Financial Times (2026). Caveat: the hallucination count originates with GPTZero, a commercial vendor; EY’s removal corroborates the core facts.
• Paul A. David, “The Dynamo and the Computer”. American Economic Review, May 1990. Electrification took roughly four decades to register in productivity figures because factories had to be redesigned around the unit-drive motor.
• OpenAI, o3 and o4-mini system card (16 April 2025); GPT-5 system card (August 2025). PersonQA hallucination rates: o1 16%, o3 33%, o4-mini 48%; OpenAI later reported significantly lower hallucination rates for GPT-5 models in browse-on and browse-off settings. Caveat: task-dependent; 1–3% on short-document summarisation per the Vectara leaderboard.
• NYC MyCity chatbot. The Markup, March 2024, documented unlawful guidance on source-of-income discrimination and workers’ tips; discontinued early 2026, with the mayor putting the cost at around half a million dollars.
• Lisanne Bainbridge, “Ironies of Automation”. Automatica, 1983.
• Endoscopist deskilling. Budzyń et al., “Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy”, The Lancet Gastroenterology & Hepatology (August 2025). Across four Polish centres and 19 experienced endoscopists, unaided adenoma detection fell from 28.4% to 22.4% after AI was introduced. Caveat: retrospective, observational design, sensitive to confounding.
• METR developer productivity RCT (July 2025). Experienced developers were 19% slower with AI tools while perceiving a 20% speedup. Caveat: 16 developers, 246 tasks, 95% CI: +2 to +39%; a February 2026 follow-up did not cleanly replicate.
• Replit autonomous agent (July 2025). The agent reportedly deleted a production database during a code freeze and fabricated about 4,000 records (Lemkin/SaaStr; Fortune; The Register; AI Incident Database #1152). Gartner (25 June 2025) separately forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027.
• McKinsey, “The State of AI” (2025). 88% adoption; 39% report enterprise EBIT impact; only 21% have redesigned any workflow; about 6% are high performers. n=1,993. Self-reported; McKinsey sells AI transformation services.
• S&P Global Market Intelligence. 451 Research, Voice of the Enterprise: AI & Machine Learning 2025 (n=1,006 IT and line-of-business professionals, North America and Europe). The share of firms abandoning most AI initiatives before production rose from 17% (2024) to 42% (2025); the average organisation scrapped 46% of proofs-of-concept. Source: S&P Global; CIO Dive (2025).
• Taco Bell drive-through AI. After deploying voice AI at 500+ drive-throughs, Taco Bell’s chief digital and technology officer said in August 2025 the company was reconsidering where to use it and would keep human order-takers at busy sites. Source: Wall Street Journal (28 August 2025). Caveat: framed by the company as iteration, not failure.
• The substrate moves. OpenAI GPT-4o sycophancy rollback, April–May 2025, with no specific deployment evaluations for the trait; the GPT-5 launch replaced multiple ChatGPT model choices and OpenAI later announced further retirements; Anthropic documentation lists rolling retirements with at least 60 days’ notice for public models.
• Practical referents. Goldman Sachs GS AI Assistant, firm-wide in 2025, multi-model. Model gateways: Portkey, Cloudflare AI Gateway, LiteLLM. Observability: Arize, Datadog LLM Observability. NatWest AI and Data Ethics Panel (a governance body). Cisco runs its internal AI platform within an eight-pillar operating model (internal; no public product page).
• Governance. FRC UK Corporate Governance Code 2024 Provision 29 applies from financial years beginning on or after 1 January 2026; sometimes compared to UK ‘SOX-lite’.
• CB Financial “shadow AI” disclosure. Community Bank (subsidiary of CB Financial Services, Nasdaq: CBFV) detected on 5 May 2026 that an employee had entered customer names, Social Security numbers and dates of birth into an unauthorised AI application; it filed an SEC Form 8-K under Item 1.05 on 11 May 2026, the first such filing attributed to shadow AI rather than an external attack. Source: SEC Form 8-K; American Banker; Wilson Sonsini (2026). Caveat: management states the data was not used to train the vendor model and that there was no material financial impact; affected-customer count undisclosed.


In our quantitative trading business, we are finding AI most powerful for accelerating the building of increasingly sophisticated automation, but the actual automation which takes actions in the business, remains deterministic rather than reasoning-based
Thanks Paul - great piece! We are using a 90s, 2000s process transformation playbook on a 2020s "alien" technology that very few people fully understand.
I wonder about the talent side of this. I love the motor analogy - we need a motor in every machine but lack the people who understand the machines and the motors that can redesign them. I'm not sure such individuals exist today - it's the FDE job but do they have true understanding of ever evolving model capabilities and quirks? I wonder if this requires a new college degree, or equivalent certification / training?