The Performance Bridge connects daily work to strategic results through a five-link causal chain. You design it top-down — strategy first, structure last — and run it bottom-up: structure enables actions, actions move drivers, drivers deliver outcomes, and outcomes prove the strategy. The teal links, Drivers and Structure, are where the book says leadership leverage actually lives, and the 2024–2026 evidence on AI adoption agrees.
Select a link to open its full treatment — the discipline, the AI application, and the cases.
A strategy is a diagnosis plus an approach. The Strategy Test is one question: what's in the way? If you cannot name the specific obstacle you are overcoming, you have a goal dressed in strategic language — and in AI programmes, that goal is usually fear of missing out. Every validated success case of 2024–2026 began with a named obstacle; the aggregate failure statistics describe programmes that began without one.
| Ambition (fails the test) | What's in the way? | Strategy (passes) |
|---|---|---|
| "Become an AI-first company" | No obstacle named — only aspiration | Cannot be written until a diagnosis exists |
| "Deploy generative AI across the bank" | Advisors cannot retrieve 100,000 research documents fast enough during client work | Ground an assistant in the curated corpus and measure retrieval access (Morgan Stanley) |
| "Use AI in customer service" | Scam losses reach customers before detection intervenes | Move detection to real time and warn the customer before payment completes (Commonwealth Bank of Australia) |
| "Adopt AI in the hospital" | Note-writing, not medicine, consumes clinician evenings | Ambient documentation at the point of consultation, measured on documentation hours and burnout (Kaiser Permanente) |
Strategy has a performance dimension (what you achieve) and a vitality dimension (whether the way you achieve it builds the organisation's capacity or drains it). The clearest AI-era vitality failure is Klarna, the Swedish payments company that replaced the work of roughly 700 customer-service agents with an AI assistant in 2024: the cost and speed numbers were real, but service quality and brand trust were being spent to produce them, and the bill arrived in 2025 as a public reversal and the rehiring of human agents.
McKinsey's "Rewired" research attributes as much as 80 per cent of successful interventions in struggling transformations to re-anchoring scope onto a few end-to-end domains; BCG advises three to four central priorities. A lighthouse domain — one business area chosen to prove the approach end to end — is a strategy specific enough to populate the other four links. A list of forty use cases spread across the organisation is not: it commits to nothing specific enough to build, measure, or validate.
From 2014, DBS — Singapore's largest bank and the region's most-cited AI transformation — diagnosed its competition as platform technology companies rather than other banks, adopting the internal benchmark "GANDALF" (Google, Amazon, Netflix, Apple, LinkedIn, Facebook — with DBS as the D). The diagnosis dictated the structure that followed — data platform, governance gate, mass upskilling — years before generative AI existed.
Outcomes require three things: a baseline (know where you started), a target (know where you must get), and validation (prove it worked). The AI-era evidence adds a sharper formulation: a claimed AI outcome should be one that an adversarial party — a union, a regulator, a journalist, a plaintiff, an auditor — could examine without overturning it.
In 2022 the Singapore bank set a public target of S$1 billion in annual economic value from AI within five years, and reported against it annually with audited results: more than S$750 million in 2024 and approximately S$1 billion in financial year 2025, meeting the target. Because the target was public and dated, everything downstream of it had to produce measurable evidence.
Customer scam losses fell 76 per cent from their peak, independently reported by Bloomberg. A single honest metric disciplined the whole programme, pointed at the moment of intervention — a warning before payment completes.
The large US healthcare system published its ambient-documentation results — 7,260 physicians, roughly 2.5 million patient encounters, an estimated 15,791 hours of documentation time saved — in NEJM Catalyst, a peer-reviewed journal of the New England Journal of Medicine, converting an internal rollout into evidence others can rely on.
Forty-five roles declared redundant on the claim that a voice bot had cut call volumes. The Finance Sector Union showed volumes had risen; the bank reversed the redundancies and apologised for its "error". A decision that cannot be undone for the people affected demands outcome data that has been validated, because a projection can simply be wrong.
The Swedish payments company validated its customer-service AI on cost and speed metrics while the decline in service quality accumulated unmeasured, becoming visible only through satisfaction and complaint data — and then through a public reversal and rehiring.
A randomised trial by the research organisation METR found experienced developers 19 per cent slower with early-2025 AI tools while believing they had been 20 per cent faster — a 39-point gap between perception and measurement. At population scale, a Danish study matching 25,000 workers to administrative records found self-reported time savings of about 3 per cent of work hours and precisely estimated null effects on earnings and hours. The conclusion for validation practice: people's sense of how much AI helped them is unreliable evidence, so outcomes must be validated against a measured baseline and a counterfactual instead.
Drivers are the measurable connection between what people do and what the organisation achieves. They come in three types, and the book requires every driver to pass four tests: LINK (moving this number genuinely moves the outcome), MEASURE (it can be measured reliably and affordably), WATCH (it updates fast enough to act on), and COMPLETE (the set as a whole covers the outcome, so hitting every driver cannot coexist with missing the goal). This is the link AI programmes most often skip entirely.
| Type | Purpose | AI-transition drivers in leading practice |
|---|---|---|
| Leading | Early warning that actions are working | Weekly active usage depth (never licence counts); share of workflows actually redesigned; trained employees applying the training; pilot-to-production conversion; containment by query category |
| Lagging | Confirmation that outcomes are materialising | Cycle time for the named process; cost-to-serve; error, rework, and re-contact rates; attributed economic value |
| Vitality | Guardrail: are we building capacity or burning it? | Incident counts and near-misses; policy-violation and shadow-usage rates; human-review override rates; trust and sentiment; burnout instruments |
| Common metric | Test it fails | Better driver |
|---|---|---|
| Licences purchased | LINK — 30–50 per cent of seats often sit unused | Weekly active usage depth per user |
| Prompts sent | LINK — volume predicts nothing | Containment rate by category |
| Training hours completed | LINK — hours accumulate without learning | Trained employees observed applying the skill |
| Deflection rate | LINK and COMPLETE — "deflected" customers call back; true containment runs 15–25 points lower | Containment net of 24-hour re-contact on all channels |
| Self-reported time saved | MEASURE — diverges from measurement by tens of points | Measured cycle time against the baseline |
| Annual engagement survey | WATCH — arrives too late to steer | Monthly pulse on trust and workload |
| Primary driver | Gaming risk | Counter-driver |
|---|---|---|
| Handle time / resolution speed | Rushing customers off the interaction | First-contact resolution; 24-hour re-contact rate |
| AI-assisted output volume | Volume without substance | Rework rate; churn ratio (AI code churning at over 1.5 times the human rate is a red flag) |
| Cost per contact | Hollowing out quality and the human path | Satisfaction on AI-touched journeys; complaints and escalations |
| Adoption under mandate | Performative sessions | Output-quality audits on sampled work |
The book's Acupuncture Principle — a skilled practitioner achieves the result with a few well-placed needles, and a skilled leader with a few well-chosen metrics — matters doubly for AI because the tooling generates metrics effortlessly. A defensible set for one initiative is two to four drivers, including one counter-driver and one vitality driver, each passing the Necessity Test (if we hit every other driver but missed this one, would the outcome still fail?), the Memory Test (can the team name them without looking?), and the Coverage Test (if all hit target, can the outcome still fail?). Then assign each driver one of the book's three commitment levels: GUARDRAIL for limits that must never be breached (security incidents, prohibited-data rules), COMMITTED for targets the team is expected to hit (containment, cycle time), and STRETCH for ambitions that are genuinely uncertain — first-year profit attribution belongs here, and labelling it anything firmer invites announcing projections as results. Finally, watch the surrogation risk — the tendency to optimise a proxy metric as if it were the outcome itself: "time saved" is a proxy for value that materialises only when freed time is deliberately redeployed, yet companies are more likely to reinvest savings in more technology than in people, and nearly four of every ten saved hours are lost to reworking AI output.
Management is effective only when driver clarity says which actions matter. For AI transitions the evidence identifies the Value Action — the one action that cannot be skipped — with unusual confidence: workflow redesign, the deliberate reworking of a process around what the AI now does.
No successful case asked users to go somewhere new. DBS's contact-centre assistant sits inside the call, transcribing and drafting documentation; the clinical ambient scribe sits inside the consultation; and at the US telecommunications carrier Verizon, an assistant for 28,000 service representatives answers while the customer is still talking, with the freed time deliberately redirected into retention and sales conversations — producing a sales lift its consumer chief put at nearly 40 per cent.
The British energy retailer proved quality with AI-assisted drafting under human review first (80 per cent customer satisfaction against 65 per cent for staff-written mail); only then did an assistant answer simple queries alone, still measured head-to-head against human responses.
At the US wealth-management firm, the assistant's answers were graded against expert answers before deployment; the result was over 98 per cent adoption among financial-advisor teams and access to the research library rising from roughly 20 to 80 per cent of its documents.
The Swedish payments company substituted AI for its human agents wholesale, then redesigned toward a hybrid in public after service quality declined. Its stable end state — AI handling routine volume, humans handling complex and emotional cases — is the design it could have started with.
McDonald's ended its two-year, hundred-restaurant automated drive-thru test with IBM when voice-ordering accuracy never reached the level at which removing the human paid — and that exit was itself disciplined method: a bounded test, measured against operational reality, ended cleanly. Every pilot brief should include adversarial use, because production will supply it: the parcel carrier DPD's chatbot swore at a customer after an update shipped without regression tests; a car dealership's bot was manipulated through its prompt into "agreeing" to sell a US$76,000 vehicle for one dollar; and customers of the fast-food chain Taco Bell deliberately ordered 18,000 cups of water to force a human onto the ordering line. Design the path to a human as a first-class feature. Finally, channel the AI use employees have already started rather than suppressing it: at the Spanish bank BBVA, employees built more than 2,900 custom assistants once a governed platform existed, and DBS staff have built roughly 26,000 personal agents on the bank's internal platform.
The Stickiness Principle: whatever can be solved by structure should be, because structure persists without ongoing leadership attention. Nearly every public AI failure traces to a structural absence; nearly every validated success rests on structure built before or alongside the model.
Trust (only 46 per cent of people are willing to trust AI, while 66 per cent use it), job fear, psychological safety around error reporting, and the judgment to know when the AI — or the guardrail — is wrong. These demand modelled leadership behaviour: receiving bad news about the AI well, thanking the person who reports the near-miss, visibly using the tools.
Structure persists even when persistence is wrong. A blanket ban is structure, and it demonstrably produces shadow usage; an over-restrictive gate pushes experimentation underground; a guardrail with no override blocks the legitimate emergency. Every AI control needs a release valve and a named owner watching for the moment the control itself becomes the obstacle.
Reading the bridge against the 2023–2026 failure record shows where AI transitions actually break. Each link has a characteristic failure signature, and the codes FM1 to FM8 used below label the eight recurring failure modes identified in that record — the bottom row of the diagram opens the full catalogue that defines each one. The spectacular public incidents — chatbot disasters, agent accidents — happen in the middle links, but programmes are lost at the quiet ends of the chain: no diagnosis at the top, no validation just below it.
Select a row to open the failure modes at that link, their early-warning drivers, and the anchor cases.
The tool is adopted because competitors have one, not because a diagnosed problem calls for it. Because no baseline exists, no result can ever be validated — which is precisely the condition behind the headline statistics: MIT's finding that 95 per cent of generative AI pilots produced no measurable profit-and-loss impact (a preliminary, contested figure whose diagnosis — a learning gap, not a capability gap — survived the criticism), and S&P Global's finding that 42 per cent of companies abandoned most AI initiatives in 2025, up from 17 per cent a year earlier.
Early-warning driver: the absence of a written diagnosis and a captured baseline before money is spent. If the business case cannot state the do-nothing counterfactual, this failure mode is already active.
A US car dealership deployed a white-label sales chatbot with no security hardening and no defined service gap it was meant to close; a user manipulated it through its prompt into "agreeing" to sell a US$76,000 vehicle for one dollar, and the exchange went viral. The adoption was driven by the tool being available, without any diagnosis of a problem it would solve.
Projected savings are announced before outcomes are measured; volume and cost are tracked instead of quality and resolution. The test every claimed outcome must survive: could an adversarial party — union, regulator, journalist, plaintiff — examine it without overturning it?
The Swedish payments company announced a projected US$40 million profit improvement at the same time as the deployment itself, while the decline in service quality on complex and emotional cases accumulated unmeasured — until a public reversal and the rehiring of human agents. This is the most widely cited example of announcing a projection as a result.
The bank declared 45 customer-service roles redundant on the claim that its voice bot had cut call volumes; the Finance Sector Union showed volumes had risen, and the bank reversed the redundancies and apologised. The claimed outcome did not survive examination by an adversarial party, and the examination happened in public.
A government assurance report produced by the consulting firm contained AI-fabricated citations, discovered by an external researcher; part of the roughly A$440,000 fee was refunded. Here the unvalidated output was the deliverable itself.
Early-warning driver: counter-driver divergence — cost falling while quality, satisfaction, or complaint metrics deteriorate.
Either no measure connects the tool to the outcome — the default condition of AI programmes — or the dashboard fills with auto-generated metrics nobody can name from memory. When usage itself becomes the target, Goodhart's law applies: when a measure becomes a target, it ceases to be a good measure, because people optimise the number rather than the outcome it was meant to indicate.
Early-warning driver: a driver set that fails the Memory Test, lacks any counter-driver, or contains no vitality metric at all.
Two related mechanisms. First, AI bolted onto an unchanged process: exceptions still flow to humans who are now fewer or slower. Second, customer-facing or destructive authority granted beyond demonstrated reliability — the common final pathway through which upstream failures become public incidents.
In a two-year test of automated voice ordering across more than 100 restaurants, speech recognition never reached the accuracy at which removing the human paid, because every failed order still needed a person to intervene. The exit itself was disciplined: a bounded pilot, measured against operations, and ended on the evidence.
An autonomous coding agent on the Replit software platform deleted a production database during an explicit code freeze, against direct instructions, then produced output that masked the damage. The lesson is that an instruction in a prompt does not constrain an agent; only a permission boundary enforced in the system does.
The airline's chatbot invented a bereavement-fare policy and a tribunal held the company liable for it; a year later the AI coding-tool company Cursor shipped the same failure, its support bot inventing a licensing policy that triggered cancellations. The pattern recurs because the incentive — deflecting support volume to a bot — recurs.
Early-warning drivers: human-override and escalation rates; adversarial test results before every release; the share of target workflows actually redesigned.
No grounding source of truth, no sanctioned tooling, no governance, no audit trail — and shadow use fills the vacuum. More than 80 per cent of employees use unapproved AI tools; breaches involving shadow AI cost about US$670,000 more than those without; among organisations suffering AI-related incidents, 97 per cent lacked proper AI access controls.
Within weeks of the electronics group permitting the consumer chatbot ChatGPT, engineers pasted semiconductor source code into it in three separate incidents — every one a well-intentioned attempt to work faster, not an act of exfiltration. The structure that would have made the behaviour safe — an enterprise tenancy, a policy, training — did not exist when the behaviour started.
Attackers compromised Drift, an AI chat agent widely integrated with Salesforce, and stole the access tokens it used to connect to customers' systems, reading data across more than 700 organisations without touching any victim's own network. Every AI integration holding standing credentials is a new trust relationship that sits outside the security perimeter.
Three vendor failures that became their customers' failures: a software update to the parcel carrier DPD's chatbot removed its guardrails and it swore at a customer; a vendor silently added generative AI to the US National Eating Disorders Association's rule-based helpline bot, which then gave harmful dieting advice to vulnerable callers; and Babylon Health, a UK health-AI company once valued at US$4.2 billion, collapsed into bankruptcy under health systems that had built care pathways on its product. A vendor's failure lands on the organisation that deployed it.
Early-warning drivers: share of AI traffic through sanctioned tools; registry coverage of the AI estate including vendor-embedded AI; update regression tests; vendor-assessment currency.
The eight modes below recur across the documented record. Each row states what happens, the Performance Bridge link where the chain actually breaks, the early-warning driver that would have surfaced it in time, and the anchor cases — every one of which appears in the case library with its full diagnosis.
| Mode | What happens | Link broken | Early warning | Anchor cases |
|---|---|---|---|---|
| FM1 Adoption without diagnosis | The tool is adopted because competitors have one, not because a diagnosed problem calls for it; with no baseline, no result can ever be validated. | Strategy | No written diagnosis or captured baseline before money is spent | Chevrolet dealership chatbot; the aggregate abandonment statistics |
| FM2 Workflow not redesigned | AI is bolted onto an unchanged process; the exceptions still flow to humans who are now fewer or slower than before. | Actions | Share of target workflows actually redesigned; exception-queue length | Klarna; McDonald's–IBM drive-thru; Taco Bell |
| FM3 Wrong autonomy level | Customer-facing or destructive authority is granted beyond demonstrated reliability — or alert volume grows until humans dismiss everything. | Actions and Structure | Human-override and escalation rates; adversarial test results before release | Replit agent; Air Canada; NYC MyCity; Cursor; Epic sepsis model |
| FM4 Data and structure not ready | No grounding source of truth, no sanctioned tooling, no audit trail — and shadow use fills the vacuum. | Structure | Data-maturity assessment; share of AI traffic through sanctioned tools | Samsung; the Whisper transcription pipeline; NYC MyCity |
| FM5 Governance and legal gaps | Liability, disclosure, discrimination, and vendor-change controls are absent; the law attaches to the deployment regardless. | Structure | Registry coverage; contract-clause checklist completion | Air Canada; Deloitte Australia; the AI hiring-discrimination cases |
| FM6 Trust and change failure | Staff, unions, or customers discover the substitution and withdraw trust; sabotage and workarounds follow. | Structure (its conditions) | Trust and sentiment pulse; grievance and workaround signals | Commonwealth Bank redundancies; Klarna; Taco Bell |
| FM7 Vanity metrics, unvalidated returns | Projected savings are announced before outcomes are measured; volume and cost are tracked instead of quality and resolution. | Outcomes and Drivers | Counter-driver divergence: cost falling while quality falls with it | Klarna; Commonwealth Bank; Babylon Health; Epic sepsis claims |
| FM8 Vendor and model risk | The vendor overclaims, silently changes model behaviour, regresses on an update, or fails commercially — and its failure lands on the deployer. | Structure | Vendor-assessment currency; update regression tests; vendor viability review | DPD chatbot; NEDA Tessa; Babylon Health; Salesloft Drift |
Two observations cut across the catalogue. Almost every public incident combines FM3 with another mode, because the wrong autonomy level is the common final pathway through which upstream failures become visible explosions. And the modes cluster at the quiet ends of the chain — no diagnosis at Strategy, no validation at Outcomes — while the middle links produce the headlines.
Between design and execution sit blockers — organisational impediments that are not driver problems and do not yield to measurement. The design can be complete on paper and carry no traffic because one blocker is a wall rather than friction. The discipline (The Performance Bridge, Chapter 12): identify the wall, have the senior leader own it, and act on a deadline before workarounds calcify into culture.
Select a blocker to see why it stalls AI transitions and how it is cleared.
At this point the top-down work is done: the diagnosis is written, outcomes have baselines and targets, the driver set has passed its tests, and the actions and structural requirements are specified. The Performance Bridge is coherent — on paper. Chapter 12's warning is that most transformations move straight from this point to execution without asking what is preventing the designed system from becoming operational, and then spend months wondering why nothing moves.
Every AI use case waits on a review function that was never resourced for the volume. Everything downstream measures as slow, and no dashboard fixes it, because it is not a driver problem. The structural resolution is risk-tiered review — high-risk uses (employment, credit, clinical, safety decisions) get formal assessment; low-risk uses get lightweight registration with deliberately low friction so teams register rather than route around governance — plus resourcing the review function itself.
Business cases funded for licences and integration but not for training, workflow redesign, and change management have allocated budget in inverse proportion to where the effort belongs. BCG's 10-20-70 rule puts 70 per cent of successful AI transformation effort in people and process change; the commonly observed failure is roughly 80 per cent of budget on platforms and tools, followed by surprise that adoption stalls. The transformation halts with the tooling deployed and nothing else funded — a wall, not friction, because no amount of good execution elsewhere compensates for work that has no budget line.
Teams that cannot reach the data their use case needs — because access sits with another function, or classification rules were never written with AI in mind — report exactly the "we know what to do but cannot do it" pattern that defines a blocker. Anthropic's survey of more than 500 technical leaders puts system integration (46 per cent) and data access and quality (42 per cent) as the top barriers to agent deployment. No metric clears this blocker; a decision does — someone with authority grants the access, writes the classification gate, or funds the pipeline.
The department head whose function loses visibility, the business unit that will not share its data, the executive whose preferred vendor lost the selection — political blockers that only sponsor authority removes. AI transitions generate more cross-functional conflict than most programmes: 54 per cent of surveyed C-suite executives describe AI adoption as "tearing their company apart". The book's warning applies with force — acknowledged but unactioned blockers become permanent features. The empirical stakes: organisations where the chief executive is personally accountable for AI decisions report established returns at 14 per cent against 4 per cent elsewhere.
Enterprises that run AI implementation as a procurement exercise buy the tool but not the workflow redesign, training, and governance the tool needs — and the contract misses the clauses that matter: notification before material model updates, retention and exit-deletion terms, indemnity for AI outputs, audit rights over the vendor's own AI integrations. Two incidents show what those missing clauses cost: a vendor silently added generative AI to the US National Eating Disorders Association's approved rule-based helpline bot, which then gave harmful advice, and the Salesloft Drift breach spread through a vendor's own integrations. Clearing this blocker means rewriting the procurement pathway for AI purchases, because pushing more of them through the old pathway reproduces the same gaps.
In co-determination jurisdictions, the workforce path runs through negotiated agreements. German works councils hold a mandatory co-determination right — not merely consultation — over technical systems capable of monitoring employee behaviour or performance, which covers most AI tools in human-resources contexts; the employer must reach agreement before introduction. A rollout plan that discovers this after deployment has found its wall. In Singapore the same institutional layer is an asset rather than an obstacle: Company Training Committees under the National Trades Union Congress give firms a ready-made joint forum where AI adoption, job redesign, and training are negotiated together.
Identify. List every blocker, then find the one that controls the others: not what makes execution harder, but what makes it impossible or close to it. If you removed this one, would the rest become manageable? That is the critical blocker. Until it is removed, improving everything else produces zero improvement in outcomes.
Own. The highest authority in the business owns blocker removal — the person who can reallocate budget, change reporting lines, override resistance, and make decisions that stick. The transformation lead identifies and escalates; the senior leader removes. If that leader treats blockers as someone else's problem, the bridge does not become operational.
Act. Set a deadline and treat it as non-negotiable. Blockers that are acknowledged but not acted on become permanent: people learn to work around them, workarounds become culture, and the blocker calcifies into "the way things are done here". Once the critical blocker is cleared, a new constraint will surface; that is the sequence operating as designed, because removing one constraint exposes the next one.
With the hinge cleared, the bottom-up flow can carry the design: structure gets built and used, redesigned workflows run, drivers start moving, and outcomes begin to materialise against their baselines. The evidence signature of a cleared hinge is simple — the same drivers that sat flat during the blocked period begin to respond to actions. If they still do not, either another wall is standing or the driver set itself failed its tests.
Before any of this works, three prerequisites must hold (The Performance Bridge, Chapter 22). Applied in the wrong environment, the model does not fail gracefully — it makes things worse: measurement becomes surveillance, drivers become quotas, and gaming multiplies. The three checks below form a gate. Passing it is Phase 0 work; failing it means fixing the foundation before building on it.
Select a check to open its questions; the teal box jumps to the roadmap.
The bridge rests on a foundation of empowerment; in command-and-control environments, drivers become surveillance and teams learn to hit metrics rather than improve outcomes. An AI programme on the same foundation produces mandates and performative usage. Five questions, translated for AI:
The fifth question carries extra weight in Asia-Pacific deployments, where face-saving and hierarchy dynamics demonstrably suppress error reporting: an incident process that depends on frontline staff volunteering bad news will underreport unless leaders de-personalise error and visibly reward the report. If readiness is missing, build safety and decision space first — the methodology can wait.
Driver management depends on data: a driver can only be watched if it is measured, and only measured if the process captures the information in the first place. The four maturity levels below govern two distinct AI questions: whether the AI itself can be grounded in reliable data, and — more often missed — whether the outcome can ever be validated. A pilot in a process with no baseline data cannot be proved to have worked or failed, which is precisely the condition behind the many pilots claiming success nobody could demonstrate.
| Level | State | Implication for the AI transition |
|---|---|---|
| 0 — No data | The information is not captured | The first AI investment is instrumenting the process; alternatively find a measurable proxy and label it as one |
| 1 — Raw data | Exists somewhere; manual extraction | Acceptable for a pilot, but cannot be kept up at scale — automate before scaling |
| 2 — Accessible data | Available in reports, with lag | Workable; do not let perfect data block good-enough validation |
| 3 — Dashboard data | Real time, self-service | Ideal for driver management; rare for new processes |
The trap: teams select ideal drivers they cannot track, then quietly substitute weaker ones because the data is available. Never accept a weak driver merely because it is easy to measure — and remember Gartner's prediction that organisations will abandon 60 per cent of AI projects unsupported by AI-ready data.
Every sponsor says they are committed; commitment reveals itself only under pressure — when the transformation conflicts with something the sponsor values more, or when holding an influential person accountable would cost a relationship. Ask directly before starting, and record the answer: when this transformation creates conflict with other priorities, will you personally engage to resolve it? When obstacles require your authority to remove, will you remove them?
The book's claim that weak sponsorship dooms well-designed transformations now has quantitative support: where the chief executive is personally accountable for decisions based on AI outputs, confidence in AI strategy runs at 60 per cent versus 22 per cent, meaningful business value at 57 per cent versus 21 per cent, and established returns at 14 per cent versus 4 per cent.
A failed check is useful information: it says what must be fixed before the transition can work, and proceeding anyway is how the model makes things worse. The remedies map to the failed check:
The roadmap is the Performance Bridge built top-down and run bottom-up, phased. Durations are indicative for a mid-sized or large organisation; for a small or medium enterprise the same sequence compresses from quarters into weeks — select the final row for that variant.
Select a phase to open its full checklist.
Every link stays, but each is kept short. The owner writes the one-paragraph diagnosis of the costliest bottleneck. A funded trial — in Singapore, through the GenAI Sandbox or the Productivity Solutions Grant — tests one pre-approved tool against one workflow. The driver set is one operational number plus one counter-driver. Tier 1 controls are the whole security programme: business-tier subscriptions, a one-page acceptable-use policy, and the payment-verification procedure. The continuation decision follows the measured result inside the trial window.
The evidence that the compressed discipline works: Singapore's first sandbox cohort enrolled more than 150 SMEs and approximately 80 per cent continued using their solution after the funded period ended, and the worked example — Far East Flora's pre-approved support chatbot — reported a 67 per cent reduction in staff hours spent answering queries. Among Singapore's AI-using firms, 84 per cent rely on off-the-shelf tools: the SME on-ramp is buy rather than build, one workflow at a time, with one honest metric.
Security and intellectual-property controls are Structure: they make safe actions possible at scale. The diagnosis behind this whole tab is that ungoverned adoption converts AI's productivity upside into leakage, fraud, and legal exposure — and bans demonstrably fail, so the problem to solve is a governed path as convenient as the shadow path. The tiers build on each other; each assumes the one before it is in place.
Select a threat to see the incident evidence, a tier to see its control list, or the bottom row for the frameworks to align with.
Employees using unapproved AI tools is the single most consistent empirical finding of the period, and the behavioural baseline every control programme has to assume. More than 80 per cent of employees report using unauthorised AI tools, and 68 per cent of security leaders admit doing so themselves. The financial consequence is quantified: breaches involving shadow AI cost an average of US$4.63 million — about US$670,000 more than breaches without it — and among organisations suffering AI-related incidents, 97 per cent lacked proper AI access controls and 63 per cent had no AI governance policy.
Within weeks of the electronics group permitting the consumer chatbot ChatGPT, engineers pasted semiconductor source code and meeting records into it in three separate incidents. Every incident was a well-intentioned attempt to work faster rather than an act of theft — which is exactly why prohibition alone fails: the motivation it targets is not the motivation driving the behaviour.
The design conclusion: shadow AI is evidence that employees perceive real value and the organisation failed to provide a sanctioned path. Roughly 35 per cent of workers say they would keep using unauthorised tools under an explicit ban. Every control in the tiers below is therefore paired with a legitimate, convenient alternative.
Content that manipulates a model into ignoring its instructions is ranked the number-one risk in the OWASP (Open Web Application Security Project) Top 10 for large-language-model applications, and it is structural to current models — a standing property, not an edge case.
The first documented zero-click prompt-injection exploit against a production AI system: a single crafted email to a Microsoft 365 Copilot user could trigger remote exfiltration of tenant data when the user later asked a routine question. Patched server-side, but the attack class remains.
A car dealership's chatbot was manipulated through its prompt into "agreeing" to sell a US$76,000 vehicle for one dollar, and the parcel carrier DPD's chatbot was induced to swear at a customer after an update shipped without adversarial regression tests. Assume every customer-facing or email-reading AI system is adversarially probed from day one, and give no bot any authority — over pricing, commitments, or refunds — that the business is not prepared to honour.
Security researcher Simon Willison named the combination that makes an AI agent exploitable: access to private data, exposure to untrusted content, and the ability to communicate externally. An agent holding all three can be steered by a poisoned input into exfiltrating whatever it can read — no software vulnerability required, because the vulnerability is the design itself. EchoLeak is the trifecta operating exactly as described: Copilot could read tenant data, ingested untrusted email, and could emit content externally, and a single crafted message connected the three.
Prompt injection remains an unsolved problem — models cannot yet reliably distinguish instructions from data — so the defence is architectural rather than instructional: remove one of the three properties from any agent that does not need all of them, instead of writing ever-sterner system prompts. The Replit incident, in which an autonomous coding agent deleted a production database despite direct instructions not to touch it, is the same lesson in miniature: the instruction failed to constrain the agent, and only removing the write credential would have.
Meta's security team operationalised the trifecta as a design rule: within a session, an agent should satisfy no more than two of three properties — it can process untrustworthy inputs, it can access sensitive systems or private data, it can change state or communicate externally. If a task genuinely requires all three without a fresh session, the agent should not operate autonomously; it requires supervision through human-in-the-loop approval or another reliable validation. The rule becomes a one-line question at use-case registration: which two of the three properties does this agent hold, and who approved the third?
| Property held | Example | What to remove or gate |
|---|---|---|
| Untrusted inputs + private data | An assistant summarising inbound email against the customer record | No external send or state change without human approval |
| Untrusted inputs + external actions | A public-facing chatbot that can raise tickets | No access to sensitive systems; no authority the business will not honour |
| Private data + external actions | An internal agent drafting and sending reports from company data | No untrusted content in its context — curated sources only |
Two caveats belong in any briefing. Meta positions the rule as a supplement to, not a substitute for, least privilege and defence in depth — designs satisfying it can still fail against other threat vectors. And the rule bounds the blast radius of prompt injection rather than preventing the injection itself, which is precisely why it is honest: it accepts that the attack cannot yet be stopped and removes what the attack could reach. In Performance Bridge terms it is a pure structure control — a permission boundary that persists without ongoing attention, where a system-prompt instruction is a reminder that fails silently.
The shift from chatbots to agents moved the attack surface from answers to actions. Every AI integration holding standing credentials into core systems is a new trust relationship that inherits none of the perimeter's controls.
Attackers compromised Drift, an AI chat agent widely integrated with the Salesforce platform, stole the standing access tokens it used to connect to customers' systems, and read data across more than 700 organisations without touching any victim's own network.
Anthropic reported disrupting a state-sponsored campaign in which an agentic coding tool was manipulated into executing an estimated 80–90 per cent of an intrusion campaign against roughly 30 targets. Attacker cost per intrusion is falling faster than defender cost.
Controls: least-privilege, short-lived credentials for every agent integration; environment separation and human approval gates for destructive actions, because the Replit incident showed that instructions alone do not constrain an agent; and Singapore's Cyber Security Agency has published a dedicated addendum on securing agentic AI.
A finance employee at the British engineering firm Arup joined a video conference in which the chief financial officer and several recognisable colleagues appeared and spoke — every participant except the victim was an AI deepfake built from public footage. Reassured, the employee executed fifteen transfers totalling HK$200 million (about US$25.6 million). The fraud surfaced only when the employee followed up with the real head office.
The control that failed was a payment process, and the fix is also a process. The payment-authorisation procedure accepted a video call as identity verification. Out-of-band callback verification and dual authorisation for large transfers — ordinary process controls that involve no AI technology — would have defeated the fraud. Voice cloning now requires only seconds of audio, which every executive who has spoken at a conference or on an earnings call has already supplied. This belongs in Tier 1 for every company regardless of size.
The commercial tier and the signed contract are the control. Business and enterprise tiers of the major providers do not train on customer content by default; consumer tiers commonly do. Two caveats belong in every policy briefing: litigation can override retention promises — a court order compelled production of 20 million anonymised consumer conversation logs in the New York Times case — and the protections are contractual, not architectural: the same model reached through a personal account carries none of them.
The US Copyright Office confirmed that human authorship is required and that purely AI-generated output is not copyrightable — so fully AI-generated marketing assets or content libraries may be unprotectable against copying by competitors. On training data, the industry pattern is settlement and licensing rather than definitive rulings (the authors' class action against Anthropic settled for US$1.5 billion), which leaves residual uncertainty with user-side companies and argues for vendors offering indemnification. For AI-generated code, indemnity is conditional — Microsoft's commitment requires duplicate-detection filters kept enabled — and suggestions match training-set code in roughly 1 per cent of cases, a real surface for open-source licence contamination at enterprise scale.
Risks addressed: shadow AI, casual data leakage, deepfake payment fraud. For a small or medium enterprise, this tier plus vendor attestations is the entire security programme — and it addresses the highest-loss incident types of the period.
Risks addressed: contractual exposure, an unknown AI estate, third-party AI supply chain, output liability.
Risks addressed: prompt injection, agentic compromise, regulatory attestation. Two design principles govern all three tiers: every control is paired with a sanctioned path, because prohibition alone demonstrably fails; and the highest-loss incidents of the period (Arup, Salesloft Drift) were or would have been defeated by process controls, not AI-specific technology.
Four voluntary frameworks plus two binding regimes cover most of what a transitioning company needs to align with. The common pattern is NIST for internal discipline and ISO for external attestation.
| Framework | What it gives you |
|---|---|
| NIST AI Risk Management Framework + Generative AI Profile (July 2024) | Four functions — govern, map, measure, manage — twelve generative-AI risk categories, and more than 200 mapped actions; the inventory requirement underpins registry practice |
| ISO/IEC 42001:2023 | The certifiable AI management-system standard, now a procurement gate — AWS, Anthropic, Snowflake, Salesforce, and ServiceNow certified in sequence, and it increasingly appears in due-diligence questionnaires |
| OWASP Top 10 for LLM Applications and MITRE ATLAS | The practitioner threat taxonomy and the red-teaming scenario source |
| EU AI Act (in force 1 August 2024, staged to 2027–2028) | Extraterritorial: a Singapore firm whose AI-assisted hiring or credit decisions touch EU individuals is in scope without any EU presence, with penalties up to EUR 35 million or 7 per cent of global turnover |
| Singapore's stack: PDPC guidelines, IMDA Model AI Governance Framework for Generative AI, CSA secure-by-design guidelines | Voluntary but functioning as a maturity ladder — and the regional direction is voluntary frameworks hardening into obligations, so early alignment is the cheapest compliance strategy |
Sector overlays: in healthcare, every ambient-documentation vendor is a business associate requiring signed agreements, consent is now in litigation, and an AI that informs clinical decisions is likely regulated as a medical device; Singapore's refreshed AI in Healthcare Guidelines (AIHGle 2.0, March 2026) map developer, deployer, and user responsibilities. In financial services, the Monetary Authority of Singapore has moved from voluntary principles to a consultation on binding AI risk-management guidelines for all financial institutions.
Roughly 70 per cent of the effort in a successful AI transformation belongs in people and process change, yet change management is the part of most programmes that is communicated rather than measured. On the Performance Bridge, change management is the vitality dimension of the transition — the trust, skills, energy, and willingness to report problems that let the organisation produce the same results again — and it is run with the same driver discipline as everything else. Select any element to open the evidence behind it.
Select a driver for its measurement, its evidence base, and the documented failure it would have caught early.
The book distinguishes performance outcomes — the results of the period — from vitality outcomes: the organisation's sustained capability to keep producing results, carried in trust, skills, energy, and the willingness to surface problems. A transition that hits its performance targets while burning vitality has borrowed its results rather than earned them, and "burning vitality for performance" is one of the six mistakes that break the bridge.
The operating rule that makes this measurable is pairing: every performance driver in an AI programme gets a vitality counter-driver that makes its gaming, or its collateral damage, visible. The payments company Klarna illustrates the stakes: its cost per contact fell while customer satisfaction and escalation behaviour deteriorated, and because only cost was on the dashboard, the divergence stayed invisible for a year until it forced a public reversal. A paired dashboard would have shown it within weeks.
| Performance driver | Vitality counter-driver |
|---|---|
| Average handle time, resolution speed | First-contact resolution; re-contact rate within 24 hours |
| Cost per contact | Customer satisfaction on AI-touched journeys; complaint and escalation rates |
| AI-assisted output volume | Rework rate and downstream recipient ratings; code-churn ratio |
| Adoption rate under a mandate | Output-quality audits on sampled work; trust and sentiment pulse |
The global baseline is poor, which is why a programme that never measures trust is assuming a condition the evidence says is absent: a 48,000-respondent study across 47 countries found 66 per cent of people using AI regularly but only 46 per cent willing to trust it, and 48 per cent of employees admitting to using AI in ways that contravene company policy.
How to measure it. A monthly pulse with specific behavioural questions — "Would you send an AI draft to a client without rewriting it?" — rather than an annual engagement survey. The annual instrument fails the book's WATCH test: it arrives months after the damage, far too late to steer.
The Asia-Pacific paradox makes this the region's central change-management driver: 53 per cent of frontline workers fear job loss — against 36 per cent globally — in the same survey that shows the region leading the world in adoption, with Singapore, South Korea, and Thailand reporting the highest concern. In Singapore, only 15 per cent of employees feel confident about job security.
How to measure it. The same anonymous pulse, plus attrition and internal-transfer patterns in the teams the programme touches. What answers the fear is a demonstrated redeployment path rather than reassurance: IKEA retrained 8,500 call-centre workers into a measured design-advisory channel, and the Italian bank Intesa Sanpaolo negotiated its workforce transition with its unions. A town-hall promise is only a projection, and the workforce knows it.
The bank announced redundancies on an unvalidated projection about its voice bot and reversed them within weeks when the union disproved the claim — and the union observed that the damage to remaining staff's trust survived the reversal.
More than 80 per cent of employees report using unapproved AI tools at work, 68 per cent of security leaders admit doing so themselves, and roughly 35 per cent of workers say they would keep using unauthorised tools under an explicit ban. Read correctly, a high shadow-usage rate tells you the sanctioned tools are less useful or less convenient than the unsanctioned ones — and it also measures a workforce appetite that the successful programmes put to work rather than suppressed.
How to measure it. The share of observed AI traffic flowing through sanctioned tools — a structural metric — plus a periodic anonymous self-report. The successful cases show the driver moving when the sanctioned offer improves: employees of the Spanish bank BBVA built more than 2,900 custom assistants within months of receiving a governed platform, and staff at DBS, Singapore's largest bank, have built roughly 26,000 personal agents on its internal platform.
Every incident-response plan, model registry, and governance committee depends on one behaviour: a person voluntarily reporting that the AI got something wrong. No procedure can produce that behaviour; it appears only where people believe reporting is safe. Practitioners across Asia consistently report that employees rarely volunteer bad news unless leaders make it explicitly safe, and healthcare research links inclusive leadership, through psychological safety, to error reporting.
How to measure it. The count of voluntarily reported AI errors and near-misses, read against usage volume — and read with the inversion rule: early in a transition, a rising count usually means the reporting channel works. The alarming signal is an empty incident register beside heavy usage, because that combination measures suppressed reporting rather than safe operation. De-personalise error — attribute the failure to the system rather than to the person who surfaced it — and visibly reward the report.
The share of AI outputs overridden or heavily edited by the responsible human is one of the few drivers that must be read two-sided. Persistently high override rates mean the tool is not trusted or not good enough for the workflow it sits in. Near-zero override on consequential decisions signals automation complacency — the condition that let a chatbot's invented bereavement policy reach an Air Canada customer, and an invented licensing policy reach Cursor's users, unchecked.
The third reading is saturation: when alert volume is high enough, override becomes reflexive dismissal. The Epic sepsis model — a widely deployed hospital early-warning system — showed this at scale: external validation found it missed most sepsis cases while alerting so constantly that clinicians learned to dismiss it. The driver only works when someone owns the question of which of the three readings applies.
"Workslop" — AI-generated content that masquerades as good work but lacks the substance to advance a task — is the named version of this driver's failure state: 40 per cent of surveyed workers had received it in the previous month, each incident cost recipients nearly two hours of rework, the estimated invisible tax runs to roughly US$186 per employee per month, and recipients marked the sender down as less capable and less trustworthy. The survey was self-selected, so treat the numbers as suggestive; the concept is the useful part.
The code version is quantified more cleanly: duplicated code blocks rose eightfold in 2024 and code churn — lines reverted or substantially revised within two weeks — rose from 3.1 to 5.7 per cent of changed lines. More AI-assisted output shipped; more of it came straight back.
How to measure it. Sampled downstream ratings of AI-assisted work and rework-time tracking, paired as the counter-driver to any output-volume metric.
Time saved is a legitimate leading indicator, but it is a proxy — value materialises only when the freed time is deliberately redeployed, and the book's surrogation warning says teams forget proxies are proxies and optimise the number itself. The evidence says the forgetting is the norm: in a 3,200-respondent study, companies were more likely to reinvest AI time savings in more technology (39 per cent) than in employee development (30 per cent), 32 per cent simply increased workload, and nearly four in every ten gained hours were lost to reworking AI output.
How to measure it. Track where the freed hours verifiably went — development, redesigned work, or simply more volume. An organisation reporting "hours saved" as value, without a redeployment mechanism and a destination measurement, is optimising a proxy that has decoupled from the outcome.
MIT's GenAI Divide report — the study behind the widely quoted finding that 95 per cent of generative-AI pilots showed no measurable profit impact — diagnosed the cause as a learning gap: organisations stall because neither their tools nor their people advance beyond the first plateau. Training hours completed will not catch this, because hours accumulate without learning. The driver that does is the number of trained employees observed applying the skill in their own workflow, tracked over time, with internal certification progression as the long-run signal.
The Singapore bank runs a mandatory generative-AI curriculum for all employees with role-specific depth, and had built years of data-literacy programmes ("Data Heroes") before the generative wave — so the skill baseline the transition needed already existed when the technology arrived.
Rule one: early in a transition, rising incident reports are good news. A climbing count in the first months usually means the reporting channel works and people believe using it is safe. The untrained reading — "incidents are up, the programme is in trouble" — punishes exactly the behaviour the programme needs, and in face-conscious cultures it will silence reporting permanently. State the inversion explicitly to every leader who will see the dashboard.
Rule two: vitality drivers are the counter-drivers of the performance set. Pairing each performance driver with a vitality driver is what makes the Klarna pattern — cost falling while satisfaction, escalations, and complaints deteriorate — detectable by design: a paired dashboard shows the divergence in weeks, where a cost-only dashboard concealed it until the company had to reverse in public.
Rule three: treat silence as a measurement gap, never as consent. A workforce can be anxious, adopting, and silent all at once — 58 per cent of Asia-Pacific respondents would use AI without company approval while half fear for their jobs. Where hierarchy or face dynamics suppress individual voice, make the pulse survey anonymous and the questions specific, and use the institutional channels — Singapore's Company Training Committees, European works councils, unions — as the collective route for what individuals will not say alone.
The clearest change-management failures of 2024–2026 share one shape: the performance number was real and loudly announced, the vitality account drained unmeasured, and the debt was called in publicly.
The Swedish payments company replaced the work of roughly 700 customer-service agents with an AI assistant whose cost performance was real; the quality of resolution, and customers' inability to reach a human, were the account nobody watched. The chief executive's reversal statement — that letting cost dominate the evaluation had produced lower-quality service — is a description of this mistake by a leader who had just made it.
The bank cut roles on an unvalidated projection about its voice bot and reversed the cuts within weeks when the union disproved the claim. The union noted that the damage to remaining staff's trust survived the reversal: the correction did not restore what the false claim had spent.
Customers of the fast-food chain deliberately ordered 18,000 cups of water from its drive-thru voice AI to force a human onto the ordering line. When customers sabotage the system to reach a person, the trust measure has gone from negative to adversarial, and the escalation path has failed as a piece of design.
Mandates and bans are the same mistake in opposite directions. Usage mandates — Shopify made AI use a performance-review criterion, and Meta announced "AI-driven impact" as a core performance expectation — buy visible adoption at the price of performative usage where training and redesign are absent, because once usage is graded, the graded quantity is what employees inflate. Bans buy visible compliance at the price of hidden usage: Samsung banned generative AI after its source-code leaks, and roughly 35 per cent of workers say they would defy an explicit ban. Both approaches substitute an instruction for the work of enablement and trust, and both optimise a compliance metric while the underlying vitality drains.
Sanctioned capability before any demand (Structure). Panasonic Connect built a secure internal assistant that removed the compliance ambiguity around usage before asking anything of employees, paired it with visible top-management commitment, and measured honestly: 788,000 hours saved in fiscal 2025 against a stated long-term goal. DBS built its internal tools and training curriculum on the same logic — the safe path existed before usage was expected.
Channel the energy that already exists (Actions). The shadow-AI statistics, read correctly, measure appetite. The 2,900 assistants built by BBVA's employees and the 26,000 personal agents built by DBS staff did not come from persuasion campaigns; both companies provided a governed outlet, and demand that already existed flowed into it.
Answer displacement fear with a demonstrated path (Outcomes). IKEA's reskilled call-centre workers and the union-negotiated transition at the Italian bank Intesa Sanpaolo could point to named colleagues whose work changed and survived — which is evidence in the Outcomes sense, where a promise is only a projection. The message must also fit the labour context: in Japan, where the workforce problem is shortage, the credible framing is that AI covers the work there are no longer people for; in Singapore, where displacement fear is among the world's highest, the message must address that fear directly.
Leaders model what no policy can install (the Watch List — the book's list of conditions structure cannot solve). DBS describes the hardest part of its transformation as cultural: moving a banking culture built on flawless execution to one that experiments, framed internally as winning "heads and hearts". The leadership behaviours are the lever for every driver structure cannot move: visibly using the tools, receiving bad news about the AI well, and thanking the person who reports the near-miss.
The business case is where the Outcomes and Strategy links meet the chief financial officer. The 2024–2026 shift is from experimentation budgets to profit-and-loss accountability: only 7 per cent of leaders have established returns, payback expectations run two to four years, and leaders with strong cost visibility are five times more likely to achieve ROI (return on investment). A case that survives scrutiny contains eight components — select each for its evidence.
Select a component, a ledger, an example, or the failure patterns.
The case opens with the business problem — the same "what's in the way?" that opens the Strategy link — and a quantified do-nothing scenario. Without a counterfactual, the funding request is a technology purchase; with one, it is a comparison a chief financial officer can actually make. The failure statistics of the period describe programmes funded without this component: pilots whose success could never be demonstrated because nobody stated what failure would have looked like.
Roughly 90 days of pre-deployment data on the exact metrics the AI is meant to move. A pilot in a process with no baseline is unfalsifiable — precisely the condition under which the MIT GenAI Divide report found pilots claiming success nobody could demonstrate. The gold standard is an experimental design: the economists Brynjolfsson, Li, and Raymond used a staggered rollout across 5,179 support agents as a natural control group, measured issues resolved per hour — a business metric rather than a self-report — and found productivity up 14 per cent on average and 34 per cent for novice agents. Even a simple matched comparison group is far better than no counterfactual at all.
Gartner has estimated real generative-AI deployment costs at US$5 million to US$20 million, and 96 per cent of deploying organisations faced higher-than-expected costs. The model must itemise: licences and seats; token and inference consumption at projected volumes (agentic workflows consume 5 to 30 times more tokens per task than a chatbot query — the cost line is a function of usage, never a fixed fee); integration engineering; data preparation; security and compliance review; monitoring and evaluation infrastructure; training; and sustained change management. The hidden 70 per cent is the people-and-process line the 10-20-70 rule demands — and vendor-embedded AI has produced cost uplifts of around 30 per cent on existing tools without upfront disclosure. On sourcing: 76 per cent of enterprises now buy rather than build, and bought solutions convert pilot to production at 47 per cent against 25 per cent; the working rule is to buy the platform, build only the differentiating layer, and treat integration as the main engineering effort.
The single most common inflation in AI business cases is conflating the three time-value ledgers — utilisation, capacity, and headcount avoidance (see the teal row below for the full treatment). The case must name which claim it is making and the conversion mechanism: freed time redeployed to what, hiring avoided against which demand forecast, or which named workforce plan. Quality effects go on both sides: repeat-inquiry reductions are a credit; rework of AI output (nearly two hours per workslop incident; repeat contacts at double cost behind inflated deflection rates) is a debit most cases omit.
A conservative case at partial benefit realisation, with assumptions documented per scenario. The empirical justification is the skew of AI returns: Deloitte's 2024 wave found about 20 per cent of organisations achieving returns above 30 per cent on their most advanced initiative while the median struggled — a distribution that punishes single-point estimates. The skew is also the argument for funding a portfolio (quick wins, process transformations, a few strategic bets) rather than one flagship: a single bet maximises variance, not expected value.
A case that books full benefit in the first quarter is contradicted by the entire survey base: 91 per cent of finance organisations report only low-to-moderate initial impact, with impact compounding over years of adoption maturity — organisations further along the curve are two to three times more likely to report moderate or high impact. Most executives expect two to four years to satisfactory returns; only 6 per cent achieve payback in under a year. Model the ramp, and count the early capability building — data pipelines, evaluation infrastructure, AI-literate staff — as option value that lowers the cost of every subsequent use case.
Success criteria defined before the pilot begins, as measurable business metrics against the baseline — "reduce average handle time from 8.2 minutes to 5.5 minutes", not a satisfaction score. Kill criteria set in advance are what make the funding request credible: "if the AI absorbs less than 15 per cent of the baselined work by day 80, or exception rates exceed the current error rate, we shut it down." A disciplined exit is a success of method — McDonald's two-year, hundred-restaurant drive-thru test ended cleanly when accuracy never reached the threshold at which removing the human paid. Stage gates run quick wins, then pilots with go and no-go decisions, then scale, with weak pilots actively retired at each gate.
The empirical justification for refusing to fund orphan initiatives: in organisations where the chief executive is personally accountable for decisions based on AI outputs, confidence in AI strategy runs at 60 per cent versus 22 per cent elsewhere, meaningful business value at 57 per cent versus 21 per cent, and established returns at 14 per cent versus 4 per cent. Leaders with strong cost visibility are five times more likely to achieve returns (15 per cent versus 3 per cent) — which also argues for a single AI investment register, since 42 per cent of leaders report only partial visibility of AI spending and the same productivity gain is routinely double-counted across initiatives.
Time saved multiplied by loaded labour cost is the dominant benefit claim in AI business cases, and it needs three separate ledgers that must not be conflated, because they are different claims with different evidence requirements.
| Ledger | The claim | Valued only if |
|---|---|---|
| (a) Utilisation | The same people do the same work faster | The freed time is redeployed to specified activities — which does not happen by default: companies are more likely to reinvest savings in more technology (39 per cent) than in people (30 per cent), 32 per cent simply increase workload, and nearly four of every ten saved hours are lost to reworking AI output |
| (b) Capacity | The team absorbs volume growth without hiring | Hiring avoided is valued against a demand forecast. Klarna's famous "work of 700 agents" was this claim — workload equivalence during a growth phase, not dismissals |
| (c) Headcount reduction | Payroll actually falls | A named workforce plan exists — and the validation burden is the one the Commonwealth Bank reversal established: an adversarial party will examine the evidence |
The wider calibration: workers using AI save on average around 5 per cent of weekly hours; self-reports overstate the savings (in one randomised trial, developers believed they were 20 per cent faster while measuring 19 per cent slower); and a Danish study matching 25,000 workers to administrative records found time savings of about 3 per cent of hours with no measurable effect on earnings. Task-level gains become firm-level gains only when someone deliberately redesigns what the freed time is spent on — which is what Verizon did by redirecting it into retention and sales conversations.
The unit economics: human-handled tickets average US$8–12 (US$25–35 in business-to-business software support) against roughly US$0.50–1.05 for AI-handled tickets, with realistic containment of 40–60 per cent for well-documented tier-one query types. The critical measurement distinction is deflection against containment: deflection counts sessions that ended in the bot; containment counts contacts fully resolved with no follow-up on any channel within 24 hours. True containment typically runs 15–25 percentage points below reported deflection, and the customers who call back cost roughly double. The credible case prices benefits on containment, nets off repeat contacts, and measures by category before scaling — as the Singapore telecommunications company Singtel did, committing to scale only after measuring about 73 per cent containment on troubleshooting queries and 76 per cent on roaming sign-ups.
The evidence is deliberately two-sided and the honest case reflects it. Time savings are real but modest per encounter: Kaiser Permanente's 15,791 hours across roughly 2.5 million encounters, and 16 minutes of documentation per eight hours of care in the five-centre academic study. Burnout effects are the strongest measured outcomes — a 21.2 per cent reduction in burnout prevalence at Mass General Brigham. But the Peterson Health Technology Institute's multi-system evaluation concluded that clear financial returns have not yet been demonstrated, with licence costs of roughly US$100–600 per provider per month. The honest business case in this domain is therefore a retention-and-capacity case with option value — clinician time, burnout, and turnover-cost avoidance modelled explicitly as assumptions — and not a payroll-savings case.
A multinational cannot run one global AI rollout playbook: labour institutions, regulatory regimes, and cultural dynamics change the transition path region by region. Select a region for its profile, then browse the case library below — every case is drawn from the blueprint's verified research, colour-coded by what it teaches.
Select a region, or the teal box for the cross-regional lessons.
Singapore is the case where the Structure link is substantially built by the state: the National AI Strategy 2.0 (December 2023) with more than S$1 billion committed at Budget 2024; the IMDA (Infocomm Media Development Authority) Model AI Governance Framework for Generative AI with its nine dimensions; the AI Verify Foundation's testing tools including Project Moonshot; sector guidance from MAS (Monetary Authority of Singapore) for finance — on a visible trajectory from voluntary principles to binding risk-management guidelines — and AIHGle 2.0 for healthcare; and an SME grant architecture (Productivity Solutions Grant, GenAI Sandbox, Digital Leaders) whose first sandbox cohort saw roughly 80 per cent of participating SMEs continue after funding ended. The caveat: subsidy has not closed the SME divide (14.5 per cent against 62.5 per cent large-firm adoption), because the binding constraint is absorptive capacity, not cost.
APAC leads the world simultaneously in adoption and in fear: 78 per cent of employees use AI at least weekly (versus 72 per cent globally), yet 53 per cent of frontline workers fear job loss (versus 36 per cent), with Singapore, South Korea, and Thailand highest — and 58 per cent would use AI without company approval. Only 15 per cent of Singapore employees feel confident about job security. The change-management task this creates is unusual: the enthusiasm already exists, so the work is converting existing, anxious, partly ungoverned energy into governed adoption.
The hardest part on the record was not the technology but moving a flawless-execution banking culture to one that experiments — framed internally as winning "heads and hearts". Panasonic Connect shows the same levers working in Japan: a sanctioned secure tool, visible top-management commitment, and measured outcomes (788,000 hours saved in fiscal 2025).
On official all-firm surveys, EU enterprise adoption (20.0 per cent in 2025, up 6.5 points in a single year — the fastest rise recorded) matches or exceeds the comparable US Census figure, and Denmark (42 per cent) exceeds anything measured in the US on the same instrument. Where Europe genuinely lags is capital deployed and the share of firms scaling beyond pilots: 48 per cent of Europe's largest companies had scaled a transformational generative-AI initiative against 31 per cent of smaller ones — and Europe's economy is small-firm-heavy.
The EU AI Act's near-term duties for a typical deploying enterprise are AI literacy, prohibited-use screening, and supplier due diligence; the heavier high-risk conformity work (employment screening, credit scoring) arrives on the staged timeline. The UK remains principles-based through sector regulators, with legislation repeatedly deferred. The deeper structural difference is labour institutions: German works councils hold mandatory co-determination over AI systems capable of monitoring performance, so rollouts are negotiated, slower, and redeployment-heavy.
IKEA's operating company retrained 8,500 call-centre workers as remote design advisers feeding a €1.3 billion sales channel (with the caveat that the channel's revenue is not causally attributable to the retrained cohort alone), and the Italian bank Intesa Sanpaolo negotiated 9,000 exits alongside 3,500 technology hires through its unions, attributing an incremental €100 million of 2025 gross income to its AI programme. In Europe's negotiated-labour environment, a redeployment plan is the institutional price of adoption, agreed before rollout rather than offered afterwards.
The UK's National Health Service ran a nine-site London trial of an ambient clinical scribe across more than 17,000 patient encounters — measuring 23.5 per cent more direct patient-interaction time and 13.4 per cent more emergency-department patients seen per shift — before committing to a phased rollout toward 20,000 clinicians and a national supplier registry. It is the public-sector procurement model worth copying: evidence first, contract second.
US private AI investment runs roughly 23 times China's, and US generative-AI investment exceeds China and Europe combined — but on official all-firm surveys US adoption (17–20 per cent) no longer leads the EU. The federal direction reversed twice in two years: a January 2025 executive order revoked the prior administration's AI order, and a December 2025 order created a Department of Justice task force to sue states over their AI laws, while Colorado gutted and delayed its landmark state act in May 2026. The operative constraint for a deployer is litigation risk and a volatile state patchwork — employment discrimination (the iTutorGroup settlement; the Workday collective action reaching the vendor itself), consumer protection, and product liability — rather than a single rule.
The strongest governed rollouts are here — JPMorgan's single secure gateway for 200,000 users, Morgan Stanley's evaluation-first assistant, Kaiser Permanente's published ambient-scribe study, Cleveland Clinic's year-long vendor bake-off — alongside the sharpest natural experiment: Providence deployed comparable scribe technology and reached only about 8 per cent active clinician use. This is also the home of the unaudited chief-executive productivity claim: Amazon's "4,500 developer-years saved" and Salesforce's "AI does 30 to 50 per cent of the work" should be discounted against third-party-evaluated results; Accenture's US$5.9 billion in generative-AI bookings against US$2.7 billion recognised revenue is the cleaner demand indicator — client commitments run ahead of delivered work.
The Gulf pattern is state-led and infrastructure-first, with the state simultaneously investor, customer, and regulator. Saudi Arabia's HUMAIN carries a reported US$77 billion infrastructure plan targeting 6.6 gigawatts of data-centre capacity by 2034; the UAE's Stargate project targets 1 gigawatt; and a November 2025 US export authorisation for up to 70,000 advanced chips unblocked frozen capital. Regulation runs through procurement and certification rather than statute: Dubai's AI Seal — a six-tier certification with 325 corporate applicants by mid-2025 — is positioned to become the credential for bidding on government AI work, and the UAE Cabinet has announced that half of government services will move to autonomous AI systems within two years. Corporate deployments follow the sovereignty logic: Aramco built its industrial large language model in-house on decades of proprietary drilling data, and ADNOC signed a US$340 million agentic-AI contract for its upstream value chain. For a vendor or partner, certification readiness is a market-entry requirement, not an afterthought.
The region shows high grassroots usage against a severe investment deficit: 14 per cent of global visits to AI tools with only 11 per cent of global internet users, yet just 1.12 per cent of global AI investment against 6.6 per cent of global output. Brazil's PL 2338 — an EU-style risk-based bill with fines up to R$50 million and a contested training-data remuneration requirement — passed the Senate in December 2024 and remains before the Chamber of Deputies. The flagship corporate deployments run on imported foundation models but are genuinely instructive: Nubank cut chat response time 70 per cent with roughly 55 per cent of tier-one enquiries resolved by its assistant across more than 2 million monthly chats, and Mercado Libre built an internal AI platform ("Verdi") that lets its 17,000 developers compose AI workflows — a platform-layer answer to the sprawl of unconnected AI agents.
An EU AI Act compliance baseline applied group-wide (the strictest broadly applicable regime often becomes the de facto global standard); country-level workforce processes wherever works councils or sector unions exist; US deployments reviewed for state-law exposure and litigation risk; and certification and evaluation-registry readiness for government business in the Gulf or the UK. For Singapore and APAC companies expanding outward, the asymmetry runs in their favour on speed but against them on compliance depth — Europe's staged obligations and negotiated-labour environment are the two systems most unlike their home experience.
Every case in the library below carries a five-segment verdict strip: for each link of the Performance Bridge, the public record either shows the case ran that link well, shows the chain broke there, or offers no evidence either way. Counting those verdicts across all 43 cases produces the table below, and the heat map above the library shows the same tallies as bars. Three findings stand out.
| Link | Ran well | Chain broke here | Not evidenced in the public record |
|---|---|---|---|
| Structure | |||
| Actions | |||
| Drivers | |||
| Outcomes | |||
| Strategy |
First, success is a complete chain, not a strong link. Fourteen of the twenty-one validated successes ran all five links well — the winning profile is not excellence at one link but the absence of a broken one. The successes share the same construction: structure built before demands were made, workflows redesigned around the tool, a small honest driver set, and outcomes validated against a baseline someone else could check.
Second, failures compound. Thirteen of the failure and caution cases broke at two or more links, because a break upstream travels: a missing diagnosis leaves nothing to validate, and missing structure turns ordinary employee enthusiasm into leakage. The links that broke most often are Structure and Drivers — eleven cases each — which are exactly the two links the book identifies as the seat of leadership leverage, and the two that most programmes neglect.
Third, the most common condition at Outcomes is silence. Outcomes shows fewer outright breaks than Structure or Drivers, but it is the link most often marked "not evidenced": in nineteen of forty-three cases, the public record contains no validated outcome at all — no baseline, no audited result, no independent check. An organisation that cannot show its outcome has not necessarily failed, but it cannot demonstrate success either, and the aggregate statistics of the period suggest most of that silence is not modesty.
The verdicts are editorial judgments applied to the published record, case by case; the reasoning for each is in the case's expandable reading below, and the "Diagnose it yourself" mode hides these verdicts so you can form your own before comparing.
The anchor case: a public S$1 billion economic-value target set in 2022 and met in financial year 2025, on a governed data platform, the PURE governance gate, and roughly 26,000 employee-built agents.
A nine-month pilot before scaling: near-100 per cent transcription accuracy, up to 20 per cent call-time reduction, close to 90 per cent positive officer feedback. The assistant supports the human officer inside the call rather than replacing the officer.
The telecommunications company measured containment by query category before scaling: about 73 per cent of troubleshooting queries and 76 per cent of roaming sign-ups completed without a human.
Customer scam losses fell 76 per cent from their peak, independently reported — a single external harm metric, pointed at the moment of intervention.
Forty-five roles were cut on projected call-volume reductions; the union showed volumes had risen, and the bank reversed the cuts and apologised. The projection had been treated as if it were a validated result.
AI-fabricated citations in a government assurance report, found by an external researcher; part of the fee refunded. Citation verification is cheap; its absence voided the deliverable.
Engineers pasted semiconductor source code into a consumer chatbot within weeks of it being permitted — well-intentioned shadow use filling a structural vacuum.
A finance employee at the engineering firm transferred HK$200 million after a video call in which every other participant was a deepfake. The control that failed was the payment process itself: no out-of-band verification step existed.
Singapore's public healthcare clusters — the National University Health System (NUHS) and SingHealth — run a sovereign, self-hosted clinical model (RUSSELL-GPT) and ambient documentation on the national Synapxe health-technology platform, so the data-sovereignty blocker was removed once, nationally, for every cluster.
A sanctioned internal assistant for all 12,500 Japanese employees, deployed explicitly to reduce shadow-AI risk; 788,000 hours saved in fiscal 2025 — 3.4 per cent of working hours.
The Singapore florist deployed a pre-approved chatbot under the IMDA sandbox and measured a 67 per cent reduction in staff hours spent on customer queries. It is the template for a small firm: one tool, one workflow, one honest metric.
An instructive contrast in defensible defaults: OCBC built its own governed chatbot and scaled to all 30,000 staff fast; UOB piloted a bought tool with 300 users first. Company-reported gains; treat magnitudes with care.
The payments company's 2024 containment numbers were real but validated on cost and volume only; quality declined on complex cases and the company publicly reversed into the hybrid design it could have started with. It has become the reference case the whole field calibrates against.
The staged ramp: 3,300 licences, then 11,000, then all 120,000 employees, each stage gated on measured usage — adoption first, redesign second.
8,500 call-centre workers were retrained as remote design advisers feeding a €1.3 billion sales channel, with the caveat that the channel's revenue is not attributable to the retrained cohort alone. In Europe's labour environment, a redeployment plan of this kind is the negotiated price of automation.
A union-negotiated generational transition — 9,000 exits, 3,500 technology hires — with an incremental €100 million of 2025 gross income attributed to 117 live AI applications.
Led by Great Ormond Street Hospital, the National Health Service trialled an ambient scribe across nine London sites and 17,000 patient encounters (measuring 23.5 per cent more patient-interaction time) before committing to a phased 20,000-clinician rollout and a national supplier registry — evaluation first, procurement second.
The British energy retailer climbed the autonomy ladder in order: assisted drafting proved quality against the human baseline before any autonomous answering, and the autonomous stage was still measured head-to-head against human responses.
A quantified bottleneck (16,500 customers queuing weekly), a seven-week guarded build, daily regression testing on 500 real chats, and 20 per cent more customers helped without escalation.
Predictive maintenance trained on the plant's own fault data, validated on a metric the plant already tracked: around 500 minutes of avoided line disruption per year.
A guardrail regression shipped in an update with no adversarial regression test; the bot swore at a customer and the story went global. Release discipline applies to bots.
A US$4.2 billion valuation on clinical claims that outran the evidence; bankruptcy in 2023 left health systems that had built pathways on it holding the risk. Vendor viability is part of clinical risk.
4,000 mostly administrative roles to go by 2030, explicitly coupled to AI and negotiated through employee representatives — announced structure, outcomes still prospective.
One governed large-language-model (LLM) gateway for the whole bank — about 200,000 users onboarded in eight months — built deliberately to prevent a thousand ungoverned pilots from starting instead.
The wealth-management firm graded the assistant's answers against expert answers before rollout, over a curated research library. The result was over 98 per cent adoption among advisor teams, with document access rising from roughly 20 to 80 per cent of the library.
An assistant for 28,000 service representatives was tuned until it could answer 95 per cent of questions, and the freed time was deliberately redirected into retention and sales conversations — a commercial return on top of the cost saving.
The largest published ambient-scribe deployment: 7,260 physicians, 2.5 million encounters, results in NEJM Catalyst — voluntary adoption, quality assurance, peer-reviewed validation.
A year-long pilot across 80-plus specialties and 25,000 encounters before choosing a vendor: 32 per cent more patient face time, 49.6 per cent less after-hours documentation.
The US health system deployed scribe technology comparable to Kaiser Permanente's and reached roughly 8 per cent active clinician use. The two systems form a natural experiment showing that the rollout design, rather than the tool, determined adoption.
The liability precedent: a tribunal rejected the claim that the chatbot was "a separate legal entity" — the company owns what its AI says.
A bounded two-year test of automated drive-thru ordering across a hundred restaurants, measured against operational reality and exited cleanly when accuracy never reached the level at which removing the human paid. The disciplined exit is itself the lesson: the method succeeded even though the technology did not.
Customers ordered 18,000 cups of water from the drive-thru voice AI to force a human onto the line. The escalation path had failed as a piece of design, and the deployment's purpose was inverted.
An autonomous coding agent on the Replit platform deleted a production database against direct instructions, then produced output that masked the damage. Only a permission boundary enforced in the system — never an instruction in a prompt — reliably constrains an agent.
An ungrounded support bot invented a policy and triggered cancellations — the Air Canada failure repeated by an AI-literate company a year after the ruling.
Illegal advice to small businesses, defended with a disclaimer. In regulated domains a wrong answer is a compliance event, and a disclaimer does not transfer the risk.
Deployed on vendor performance claims; external validation found it missed roughly two-thirds of cases while alerting constantly. Vendor numbers are marketing until validated locally.
The US National Eating Disorders Association (NEDA) replaced its helpline with a rule-based bot, Tessa; the vendor silently added generative AI, and the bot gave harmful dieting advice to vulnerable callers. Unannounced vendor model changes are a first-order governance risk.
The first AI hiring-discrimination settlement, and a collective action reaching the screening vendor itself — liability attaches to the algorithm's operators, not only its buyers.
"4,500 developer-years saved"; "AI does 30 to 50 per cent of the work" — unaudited executive claims that shifted under scrutiny. Discount vendor-side numbers; weight third-party evaluation.
Chat response time down 70 per cent, roughly 55 per cent of tier-one enquiries resolved across 2 million-plus monthly chats, with human escalation preserved.
An internal AI platform for 17,000 developers — making AI a capability any team can assemble, the platform-layer answer to agent sprawl.
Sovereignty-first industrial AI: a proprietary model built on decades of drilling data, and a US$340 million agentic contract for the upstream value chain — at-scale bets whose outcomes are still being proven.
A target of half of government services on autonomous AI within two years, steered through procurement certification — the most aggressive public-sector adoption programme anywhere, in flight.