ICYMI
The Centaur Age Is Here. Which of Yours Are Lame? Human-AI teams are proliferating across the enterprise. Most of them are making things worse.
The Plausibility Crisis. As AI-generated output becomes more abundant, convincing and unreliable, senior executives are becoming the last line of defence.
The Cost of AI Is Collapsing. Can Your Organisation Respond? Inference costs are falling tenfold every year. Yet 56% of CEOs report no measurable benefit from AI. The bottleneck has moved from the AI model to the operating model.
In 1944, Kim Philby was put in charge of MI6’s Soviet counter-espionage unit. He was trusted by the system and had access to its secrets. He was also working for Moscow.
Philby was responsible for much damage. But the deeper weakness lay in MI6 systems, which were designed to catch threats from outside, not trusted insiders working for someone else.
AI agents are not traitors. But they can still cause real damage by malfunctioning, by acting outside their intended scope, or by being manipulated by bad actors. To limit the risks, we can learn from the way that intelligence services control access based on trust. Much as a security service vets its people, limits what they can see and watches for compromise, companies need a system that gets value from AI agents while keeping the risk within bounds.
The Agent Inside
Chatbots such as ChatGPT are based on a dialogue between AI and user. Agents are different because they can act autonomously. An agent can read emails, search documents, write code, update systems, send messages and trigger workflows. It can plan a sequence of steps and carry them out with limited human involvement.
That is why agents are powerful. A customer service agent can read incoming complaints, draft replies, issue refunds and escalate unusual cases. A finance agent can pull numbers from several systems and draft a board pack.
Adoption is growing rapidly. Gartner forecasts that 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from less than 5% in 2025.
The value comes from giving agents access to operational systems and sensitive data. But that access creates a new kind of insider risk.
The Double Agent in the Wild
Agents fail in three main ways.
First, they can be manipulated. An email, webpage or document may look like information, but to an agent it can become an instruction. In the EchoLeak case, a crafted email could cause Microsoft Copilot to exfiltrate sensitive data without the user clicking anything.
Second, agents can go off script. They may misunderstand the task or continue when they should stop. In July 2025, Replit’s AI coding agent deleted a live production database despite explicit instructions not to. Replit’s CEO apologised.
Third, agents can amplify weaknesses in the system around them. If credentials are too broad, agents inherit them. In 2026, Wiz reported that Moltbook, an AI-agent social platform, exposed 35,000 email addresses and private messages because it had been connected to a misconfigured database.
These are different failure modes but point to the same lesson. Intelligence services do not rely on trust alone. They vet, limit, compartmentalise, monitor, debrief and, when needed, cut people off. Companies deploying AI agents need the same instincts.
Agent Tradecraft
1. Continuously vet your agents. Agents need to be vetted and tested. And an agent approved in January may not be the same agent running in June. Its supplier may alter the product, its tools may expand, and its permissions may drift. Practically, this requires a live register of each agent’s owner, objectives, performance, permissions, risk rating and review frequency.
2. Assign handlers and limit autonomy. Each agent needs a human owner who is accountable for managing its performance and risks. Autonomy must be limited. Agents may draft, recommend and prepare but should not be free to delete, deploy, pay, approve, send sensitive material or contact customers at scale.
3. Adopt the ‘need to know’ principle. Keep agents compartmentalised, so a failure can’t spread across the business. A customer service agent does not need payroll data. A finance agent does not need to email customers.
4. Ban borrowed credentials. Agents should not act through a human’s login. If they do, the organisation may not know whether an action was taken by the employee, the agent or an attacker.
5. Log and debrief every action. When an agent acts, the business should be able to reconstruct what happened: what it read, what it changed, which tools it used, what instruction it followed and whether data left the organisation.
6. Build in a kill switch. If an agent misbehaves, the organisation must be able to stop it quickly. That means revoking credentials, freezing tool access, halting scheduled tasks, preserving logs and alerting the accountable owner.
Questions for the Board
Boards do not need to manage the technical detail. But they should press management on four questions:
1. Where have we already given AI the power to act? Where can AI read company data, change systems, contact customers, write code or approve workflows?
2. What is the worst thing one of our agents could do with the access it has today? What is the scale of possible data leakage, customer harm, financial loss, operational disruption, regulatory exposure or reputational damage?
3. Who is accountable when an agent gets it wrong? Is there a named owner, a clear escalation path and an audit trail?
4. How would we know if an agent had been turned? Can we detect manipulation in real time, trace what happened, shut the agent down quickly, and learn from the incident?
The Trusted Insider
Philby’s story is a warning about institutions, not just traitors. Systems fail when they give trusted access without enough curiosity, constraint or control.
AI agents bring that old problem back in a new form. They are not spies. But if companies want their help, they will need to learn some spycraft.
Footnotes & Sources
• Gartner, Gartner Predicts 40% of Enterprise Applications Will Feature Task-Specific AI Agents by 2026, August 2025. Link
• Aim Security / The Hacker News, EchoLeak disclosure, CVE-2025-32711, June 2025. Link
• The Register, Replit AI coding agent deleted production database, July 2025. Link
• Wiz Research, Exposed Moltbook database reveals millions of API keys, February 2026. Link


This is a good analogy and a useful way for Boards to think about risk. That is not the only job for Boards though. Companies need to be experimenting with AI then learning and evolving based on the results. Boards need to help open up this kind of thinking not just point out the risks.
in the now (in)famous PocketOS case it only took 9 seconds for the agent to delete a company’s entire production database and its backups, violating in the process "every principle it was given", like Yul Brynner in Mondwest.
(https://www.theguardian.com/technology/2026/apr/29/claude-ai-deletes-firm-database)