In June 2005, two amateurs entered the PAL/CSS Freestyle Chess Tournament with three off-the-shelf PCs and beat a field that included grandmaster-computer teams and high-end chess machines. Steven Cramton and Zackary Stephen had found a better division of labour between human pattern-recognition and machine calculation. Garry Kasparov, who launched advanced, or “centaur”, chess in 1998, concluded that the advantage came from process, not human skill or machine strength alone.
That era is now over in chess. World No. 1 Magnus Carlsen has said he has no realistic chance against a chess engine on his phone. But centaur models remain highly relevant to business. While chess is closed and has perfect information, enterprise work has ambiguous goals, hidden information and judgement. Well-designed centaurs can outperform unaided humans (and unaided machines) in many business domains. Poorly designed, or lame, centaurs can underperform both. They may look like collaborative teams, but too often they deliver polished nonsense or worse.
As in chess, the advantage comes (or doesn’t come) from process, not from human skill or machine strength alone.
The Centaur at Work
The evidence is mounting that well-designed centaurs can meaningfully impact performance. The BCG-Harvard “jagged frontier” study put 758 BCG consultants through a set of realistic consulting tasks. On tasks inside the model’s frontier, consultants collaborating with GPT-4 completed 12% more tasks, worked 25% faster and produced 40% higher-quality work. But on a separate task deliberately chosen to sit beyond the model’s capability, consultant performance was 19 percentage points worse than the unaided control group.
McKinsey’s 2025 State of AI survey of 1,993 respondents in 105 nations shows the value of centaur design. Some 88% of organisations regularly use AI in at least one function; about a third are scaling; and roughly 6% are high performers. Those high performers are 2.8 times more likely to have redesigned workflows and nearly three times more likely to have defined human-in-the-loop validation.
The Lame Centaur in the Wild
The risks are already visible. Some 51% of organisations report at least one negative AI incident in the past year. Commonwealth Bank of Australia is one example. In July 2025 it announced 45 redundancies in customer service, citing AI deflection of about 2,000 calls a week. Within weeks, the bank reversed the decision, rehired staff and apologised, saying the original assessment missed relevant business considerations. Call volumes had risen and overtime was spiking. The AI’s deflection rate had not been validated before the human role was removed.
McDonald’s and Taco Bell have made similar errors. McDonald’s ended its drive-thru voice-ordering pilot at 100+ restaurants in June 2024 after viral failures including runaway repeat-item orders. Taco Bell put voice AI into 500+ locations and processed more than two million orders, then told the Wall Street Journal in August 2025 it was rethinking. Gartner predicts that half of organisations expecting to significantly reduce customer-service workforces will abandon those plans by 2027.
Business is not Chess
Two counter-arguments are commonly cited.
The first is the chess-pattern argument: the centaur is a transitional form and AI will eventually surpass humans in enterprise work too. In some narrow tasks with closed rules and clean feedback, such as fraud detection, that is already true. But high-value activities such as M&A, regulatory negotiation and brand management are not chess. The evidence shows that the firms getting the best results are not the ones that have removed humans, but the ones that have redesigned the human role around what the machine cannot do. The chess endgame does not generalise.
The second is the cost argument: humans are too expensive to keep in the loop. The BCG-Harvard study refutes this. Outside the frontier, consultants paired with AI performed worse. The cost of pulling humans out of the loop is not zero. It includes errors that are not caught, hallucinations that become contractual terms and misjudgements that reach regulators. Net AI Yield (NAY) should reflect the full picture.
Building Centaurs that Run
Four practical actions can set your centaurs up for success:
1. Map the frontier task by task. Every workstream sits somewhere on the jagged frontier of AI capability: clearly inside it (delegate to AI, with sampled verification); clearly outside it (humans own, with AI as a sounding board); or on the boundary (use genuine collaboration with an explicit handoff protocol). The frontier moves with each model release, so revisit task allocation frequently and adjust the level of automation.
2. Redesign end-to-end processes, not just individual tasks. McKinsey’s data demonstrates that workflow redesign is the largest single predictor of EBIT impact from AI. Yet most organisations bolt AI onto an existing process and call it transformation. CBA did not fail because it chose the wrong model; it failed because it removed humans from a customer-service process without re-engineering the process around the new division of labour. The 6% of high performers in the McKinsey survey have fundamentally redesigned workflows by rethinking handoffs, escalation triggers and feedback loops, rather than simply automating steps in the old sequence.
3. Reshape roles around the new division of labour. Once processes are redesigned, roles must follow. Three shifts are central: reducing roles in transactional middle layers; expanding roles in judgement-dense work, such as client relationships, exception handling, ethical adjudication and deal pricing under uncertainty; and establishing entirely new roles such as agent orchestrators and hybrid managers of human-plus-agent teams. Allen & Overy built its Markets Innovation Group of lawyers, engineers and technologists before it deployed Harvey, for example. Less effective competitors issued copilot licences but didn’t change job descriptions.
4. Assign decision rights and escalation paths. As agents proliferate, accountability can blur. Three questions need clear answers: Who has authority to deploy an agent into a process? Who owns its output? What is the escalation path when the agent disagrees with the human, or vice versa?
Questions for the Board
Boards can steer centaur development by asking the right questions:
1. Have we modelled scenarios for the pace of AI adoption, and the consequent impact on our workforce plan, including the case where adoption is slower or more expensive than the headlines suggest?
2. Are we measuring what predicts success (e.g., cycle-time reduction, override rate, Net AI Yield, negative AI incidents) or are we measuring inputs (e.g., licences issued, prompts per head)?
3. For each process where we have deployed AI, has the workflow been redesigned end to end — or have we automated steps in the existing sequence and called it transformation?
4. When an AI agent goes wrong, who is accountable, and have they signed off on the decision rights that put them there?
5. If we deploy a faulty AI agent, how quickly can we detect it, and how quickly can we roll it back?
The Sharpest Process Wins
Cramton and Stephen did not win the 2005 Freestyle tournament because they had the best hardware or the deepest chess knowledge. They won because they had designed the sharpest process for deciding when the human should lead and when the machine should. Twenty years later, that lesson is relevant to every enterprise. It’s likely that many of these centaurs are already inside your organisation, and that some of them are limping.
The question is not whether to pair humans with machines. It is whether you have designed the process well enough that the pairing is effective.
Footnotes & Sources
• ChessBase, Dark horse ZackS wins Freestyle Chess Tournament, June 2005. Amateurs Steven Cramton and Zackary Stephen, playing as “ZackS” with three off-the-shelf PCs running Fritz, Shredder and Junior, won the inaugural PAL/CSS Freestyle event over a field including grandmaster + computer teams and powerful chess machines.
• Garry Kasparov, The Chess Master and the Computer, The New York Review of Books, February 2010. Kasparov launched “Advanced Chess” against Veselin Topalov in León in 1998 and drew the centaur lesson from the 2005 Freestyle: “weak human + machine + better process” beats both a stronger computer alone and a stronger human + machine + worse process.
• Tyler Cowen, “Centaur chess” is now run by computers, Marginal Revolution, February 2024. Source of the quotation: “for years now, the engines have been so strong that strategy no longer made sense.” See also Cowen, Average Is Over, Dutton, 2013, for the original centaur thesis.
• The Joe Rogan Experience #2275 — Magnus Carlsen, February 2025. Asked by Rogan whether he could beat his phone at the highest level, the world No. 1 replies “no, no chance.”
• Dell’Acqua, McFowland, Mollick, Lifshitz-Assaf, Kellogg, Rajendran, Krayer, Candelon & Lakhani, Navigating the Jagged Technological Frontier, Harvard Business School Working Paper 24-013, September 2023. Pre-registered field experiment with 758 BCG consultants across realistic consulting tasks. Inside the frontier: 12.2% more tasks completed, 25.1% faster, 40% higher quality; below-average performers improved 43%. On a separate task outside the frontier, AI users were 19 percentage points more likely than the unaided control to be wrong.
• Allen & Overy, A&O announces exclusive launch partnership with Harvey (press release), February 2023; A&O Shearman / Harvey, Customer Story (harvey.ai). Harvey deployed to 3,500+ A&O lawyers across 43 offices; ContractMatrix saves around seven hours from the average contract review, and efficiency gain of about 30%.
• McKinsey & Company (QuantumBlack), The state of AI in 2025: Agents, innovation, and transformation, November 2025. Survey fielded 25 June – 29 July 2025 (n=1,993 respondents in 105 nations). 88% report regular AI use in at least one business function; ~one-third are scaling AI; ~6% qualify as AI high performers; high performers are 2.8× more likely to have fundamentally redesigned workflows and nearly 3× more likely to have defined human-in-the-loop validation; 51% of all organisations report at least one negative AI-related incident in the prior year.
• Bloomberg / Finance Sector Union of Australia, August 2025. CBA reversed 45 customer-service redundancies after the FSU escalated to the Fair Work Commission with evidence of rising call volumes, manager redeployment to phones, and overtime spikes. CBA spokesperson: the original assessment “did not adequately consider all relevant business considerations and this error meant the roles were not redundant.”
• Klarna, Klarna AI assistant handles two-thirds of customer service chats in its first month (press release), February 2024; Bloomberg, Klarna Turns From AI to Real Person Customer Service, May 2025. Initial claim: assistant doing the work of 700 full-time agents. CEO Sebastian Siemiatkowski’s later reversal: “cost has unfortunately seemed to be a too predominant evaluation factor; what you end up having is lower quality.”
• Restaurant Business, McDonald’s is ending its drive-thru AI test, June 2024. McDonald’s wound down its IBM-built Automated Order Taking pilot at 100+ US locations after viral failures including runaway repeat-item orders. The chain stated it would “make an informed decision on a future voice-ordering solution by the end of the year.”
• Wall Street Journal, Taco Bell Rethinks Future of Voice AI at the Drive-Through, August 2025. After deploying voice AI at 500+ Taco Bell locations and processing more than two million orders, the chain began rethinking the rollout. Taco Bell is shifting toward voice AI off-peak with humans monitoring and stepping in at peak.
• Gartner, Gartner Predicts 50% of Organizations Will Abandon Plans to Reduce Customer Service Workforce Due to AI (press release), June 2025. By 2027, half of organisations expecting to “significantly reduce” customer-service workforces will abandon those plans (March 2025 poll of 163 customer-service leaders, 95% of whom plan to retain human agents).
• McKinsey Podcast, Trust in the age of agents (interview with Rich Isenberg), March 2026. Source of the framing: “Agency isn’t a feature; it’s a transfer of decision rights.”
• Bureau d’Enquêtes et d’Analyses (BEA), Final Report on the accident on 1st June 2009 to flight AF 447, July 2012. 25 safety recommendations covering manual handling at high altitude, automation-paradox awareness, stall-recovery training and crew coordination/CRM under startle and surprise.
• NASA, Resource Management on the Flightdeck, Conference Publication 2120, 1979. The 1979 NASA workshop, convened in the wake of the 1977 Tenerife runway collision and other accidents, is the canonical origin of Crew Resource Management — formalising communication, challenge-and-response, and crew-coordination protocols; codified for U.S. carriers in FAA AC 120-51E, 2004.

