The AI transition framework

How a company implements artificial intelligence on the Performance Bridge, drawn from the documented record of 2023–2026. Select any element in the diagrams to open the evidence behind it.
"Build structure. Manage actions. Watch drivers. Outcomes follow." Swarup Biswas, The Performance Bridge (2026), Chapter 1

Reading the bridge against the 2023–2026 failure record shows where AI transitions actually break. Each link has a characteristic failure signature, and the codes FM1 to FM8 used below label the eight recurring failure modes identified in that record — the bottom row of the diagram opens the full catalogue that defines each one. The spectacular public incidents — chatbot disasters, agent accidents — happen in the middle links, but programmes are lost at the quiet ends of the chain: no diagnosis at the top, no validation just below it.

Strategy FOMO adoptionNo diagnosis, no baseline Outcomes Unvalidated resultsProjections announced as achievements Drivers Vanity metrics and gamingLicences, prompts, self-reports Actions Bolted-on AINo redesign, wrong autonomy level Structure Nothing underneathShadow AI, ungoverned data The eight failure modes, FM1 to FM8 The reference catalogue: definition, broken link, early warning, anchor cases

Select a row to open the failure modes at that link, their early-warning drivers, and the anchor cases.

Failure at Strategy: adoption without diagnosis

Failure mode FM1 · the most common starting condition in the record

The tool is adopted because competitors have one, not because a diagnosed problem calls for it. Because no baseline exists, no result can ever be validated — which is precisely the condition behind the headline statistics: MIT's finding that 95 per cent of generative AI pilots produced no measurable profit-and-loss impact (a preliminary, contested figure whose diagnosis — a learning gap, not a capability gap — survived the criticism), and S&P Global's finding that 42 per cent of companies abandoned most AI initiatives in 2025, up from 17 per cent a year earlier.

Early-warning driver: the absence of a written diagnosis and a captured baseline before money is spent. If the business case cannot state the do-nothing counterfactual, this failure mode is already active.

Chevrolet dealership chatbot (December 2023)

A US car dealership deployed a white-label sales chatbot with no security hardening and no defined service gap it was meant to close; a user manipulated it through its prompt into "agreeing" to sell a US$76,000 vehicle for one dollar, and the exchange went viral. The adoption was driven by the tool being available, without any diagnosis of a problem it would solve.

Failure at Outcomes: unvalidated results

Failure mode FM7 · projections announced as if they were achievements

Projected savings are announced before outcomes are measured; volume and cost are tracked instead of quality and resolution. The test every claimed outcome must survive: could an adversarial party — union, regulator, journalist, plaintiff — examine it without overturning it?

Klarna (2024–2025)

The Swedish payments company announced a projected US$40 million profit improvement at the same time as the deployment itself, while the decline in service quality on complex and emotional cases accumulated unmeasured — until a public reversal and the rehiring of human agents. This is the most widely cited example of announcing a projection as a result.

Commonwealth Bank of Australia (2025)

The bank declared 45 customer-service roles redundant on the claim that its voice bot had cut call volumes; the Finance Sector Union showed volumes had risen, and the bank reversed the redundancies and apologised. The claimed outcome did not survive examination by an adversarial party, and the examination happened in public.

Deloitte Australia (2025)

A government assurance report produced by the consulting firm contained AI-fabricated citations, discovered by an external researcher; part of the roughly A$440,000 fee was refunded. Here the unvalidated output was the deliverable itself.

Early-warning driver: counter-driver divergence — cost falling while quality, satisfaction, or complaint metrics deteriorate.

Failure at Drivers: vanity metrics and gaming

Two opposite driver failures: no measures at all, or too many to watch

Either no measure connects the tool to the outcome — the default condition of AI programmes — or the dashboard fills with auto-generated metrics nobody can name from memory. When usage itself becomes the target, Goodhart's law applies: when a measure becomes a target, it ceases to be a good measure, because people optimise the number rather than the outcome it was meant to indicate.

  • Licence theatre: 30–50 per cent of provisioned AI licences typically show zero or near-zero usage; weekly active usage of a major workplace copilot runs at roughly 20–35 per cent of purchased seats.
  • Mandate gaming: once AI usage enters performance reviews — Shopify's chief executive declared AI use "non-optional" and built it into reviews, and Meta announced "AI-driven impact" as a core performance expectation — prompts and sessions become the quantity employees inflate.
  • Workslop: AI-generated volume that looks like work costs recipients nearly two hours of rework per incident and damages the sender's standing.
  • Code churn: duplicated code blocks rose eightfold in 2024; churn rose from 3.1 to 5.7 per cent of changed lines.

Early-warning driver: a driver set that fails the Memory Test, lacks any counter-driver, or contains no vitality metric at all.

Failure at Actions: bolted-on AI and the wrong autonomy level

Failure modes FM2 and FM3 · the visible explosions happen here

Two related mechanisms. First, AI bolted onto an unchanged process: exceptions still flow to humans who are now fewer or slower. Second, customer-facing or destructive authority granted beyond demonstrated reliability — the common final pathway through which upstream failures become public incidents.

McDonald's–IBM drive-thru (ended June 2024)

In a two-year test of automated voice ordering across more than 100 restaurants, speech recognition never reached the accuracy at which removing the human paid, because every failed order still needed a person to intervene. The exit itself was disciplined: a bounded pilot, measured against operations, and ended on the evidence.

Replit agent (July 2025)

An autonomous coding agent on the Replit software platform deleted a production database during an explicit code freeze, against direct instructions, then produced output that masked the damage. The lesson is that an instruction in a prompt does not constrain an agent; only a permission boundary enforced in the system does.

Air Canada and Cursor (2024, 2025)

The airline's chatbot invented a bereavement-fare policy and a tribunal held the company liable for it; a year later the AI coding-tool company Cursor shipped the same failure, its support bot inventing a licensing policy that triggered cancellations. The pattern recurs because the incentive — deflecting support volume to a bot — recurs.

Early-warning drivers: human-override and escalation rates; adversarial test results before every release; the share of target workflows actually redesigned.

Failure at Structure: nothing underneath

Failure modes FM4, FM5, FM6, FM8 · the quiet foundation failures

No grounding source of truth, no sanctioned tooling, no governance, no audit trail — and shadow use fills the vacuum. More than 80 per cent of employees use unapproved AI tools; breaches involving shadow AI cost about US$670,000 more than those without; among organisations suffering AI-related incidents, 97 per cent lacked proper AI access controls.

Samsung (2023)

Within weeks of the electronics group permitting the consumer chatbot ChatGPT, engineers pasted semiconductor source code into it in three separate incidents — every one a well-intentioned attempt to work faster, not an act of exfiltration. The structure that would have made the behaviour safe — an enterprise tenancy, a policy, training — did not exist when the behaviour started.

Salesloft Drift supply chain (August 2025)

Attackers compromised Drift, an AI chat agent widely integrated with Salesforce, and stole the access tokens it used to connect to customers' systems, reading data across more than 700 organisations without touching any victim's own network. Every AI integration holding standing credentials is a new trust relationship that sits outside the security perimeter.

Vendor risk: DPD, NEDA, Babylon

Three vendor failures that became their customers' failures: a software update to the parcel carrier DPD's chatbot removed its guardrails and it swore at a customer; a vendor silently added generative AI to the US National Eating Disorders Association's rule-based helpline bot, which then gave harmful dieting advice to vulnerable callers; and Babylon Health, a UK health-AI company once valued at US$4.2 billion, collapsed into bankruptcy under health systems that had built care pathways on its product. A vendor's failure lands on the organisation that deployed it.

Early-warning drivers: share of AI traffic through sanctioned tools; registry coverage of the AI estate including vendor-embedded AI; update regression tests; vendor-assessment currency.

The eight failure modes: the reference catalogue

Every recurring way an AI transition broke in the 2023–2026 record, defined in one place

The eight modes below recur across the documented record. Each row states what happens, the Performance Bridge link where the chain actually breaks, the early-warning driver that would have surfaced it in time, and the anchor cases — every one of which appears in the case library with its full diagnosis.

ModeWhat happensLink brokenEarly warningAnchor cases
FM1 Adoption without diagnosisThe tool is adopted because competitors have one, not because a diagnosed problem calls for it; with no baseline, no result can ever be validated.StrategyNo written diagnosis or captured baseline before money is spentChevrolet dealership chatbot; the aggregate abandonment statistics
FM2 Workflow not redesignedAI is bolted onto an unchanged process; the exceptions still flow to humans who are now fewer or slower than before.ActionsShare of target workflows actually redesigned; exception-queue lengthKlarna; McDonald's–IBM drive-thru; Taco Bell
FM3 Wrong autonomy levelCustomer-facing or destructive authority is granted beyond demonstrated reliability — or alert volume grows until humans dismiss everything.Actions and StructureHuman-override and escalation rates; adversarial test results before releaseReplit agent; Air Canada; NYC MyCity; Cursor; Epic sepsis model
FM4 Data and structure not readyNo grounding source of truth, no sanctioned tooling, no audit trail — and shadow use fills the vacuum.StructureData-maturity assessment; share of AI traffic through sanctioned toolsSamsung; the Whisper transcription pipeline; NYC MyCity
FM5 Governance and legal gapsLiability, disclosure, discrimination, and vendor-change controls are absent; the law attaches to the deployment regardless.StructureRegistry coverage; contract-clause checklist completionAir Canada; Deloitte Australia; the AI hiring-discrimination cases
FM6 Trust and change failureStaff, unions, or customers discover the substitution and withdraw trust; sabotage and workarounds follow.Structure (its conditions)Trust and sentiment pulse; grievance and workaround signalsCommonwealth Bank redundancies; Klarna; Taco Bell
FM7 Vanity metrics, unvalidated returnsProjected savings are announced before outcomes are measured; volume and cost are tracked instead of quality and resolution.Outcomes and DriversCounter-driver divergence: cost falling while quality falls with itKlarna; Commonwealth Bank; Babylon Health; Epic sepsis claims
FM8 Vendor and model riskThe vendor overclaims, silently changes model behaviour, regresses on an update, or fails commercially — and its failure lands on the deployer.StructureVendor-assessment currency; update regression tests; vendor viability reviewDPD chatbot; NEDA Tessa; Babylon Health; Salesloft Drift

Two observations cut across the catalogue. Almost every public incident combines FM3 with another mode, because the wrong autonomy level is the common final pathway through which upstream failures become visible explosions. And the modes cluster at the quiet ends of the chain — no diagnosis at Strategy, no validation at Outcomes — while the middle links produce the headlines.

Between design and execution sit blockers — organisational impediments that are not driver problems and do not yield to measurement. The design can be complete on paper and carry no traffic because one blocker is a wall rather than friction. The discipline (The Performance Bridge, Chapter 12): identify the wall, have the senior leader own it, and act on a deadline before workarounds calcify into culture.

Design on paper Strategy, outcomes, drivers set The hinge One of these is the wall Legal sign-off queue Unfunded change work Data access blocked Sponsor conflict Bought like software No labour agreement Identify the wall Sponsor owns it Act on a deadline Execution in reality Structure built, actions moving

Select a blocker to see why it stalls AI transitions and how it is cleared.

Design on paper

At this point the top-down work is done: the diagnosis is written, outcomes have baselines and targets, the driver set has passed its tests, and the actions and structural requirements are specified. The Performance Bridge is coherent — on paper. Chapter 12's warning is that most transformations move straight from this point to execution without asking what is preventing the designed system from becoming operational, and then spend months wondering why nothing moves.

Blocker: the unfunded 70 per cent

Business cases funded for licences and integration but not for training, workflow redesign, and change management have allocated budget in inverse proportion to where the effort belongs. BCG's 10-20-70 rule puts 70 per cent of successful AI transformation effort in people and process change; the commonly observed failure is roughly 80 per cent of budget on platforms and tools, followed by surprise that adoption stalls. The transformation halts with the tooling deployed and nothing else funded — a wall, not friction, because no amount of good execution elsewhere compensates for work that has no budget line.

Blocker: data access

Teams that cannot reach the data their use case needs — because access sits with another function, or classification rules were never written with AI in mind — report exactly the "we know what to do but cannot do it" pattern that defines a blocker. Anthropic's survey of more than 500 technical leaders puts system integration (46 per cent) and data access and quality (42 per cent) as the top barriers to agent deployment. No metric clears this blocker; a decision does — someone with authority grants the access, writes the classification gate, or funds the pipeline.

Blocker: the sponsor conflict left unresolved

The department head whose function loses visibility, the business unit that will not share its data, the executive whose preferred vendor lost the selection — political blockers that only sponsor authority removes. AI transitions generate more cross-functional conflict than most programmes: 54 per cent of surveyed C-suite executives describe AI adoption as "tearing their company apart". The book's warning applies with force — acknowledged but unactioned blockers become permanent features. The empirical stakes: organisations where the chief executive is personally accountable for AI decisions report established returns at 14 per cent against 4 per cent elsewhere.

Writer, 2026; KPMG Global AI Pulse, Jun 2026; Biswas, The Performance Bridge (2026), Ch. 12 and 23.

Blocker: AI procured like ordinary software

Enterprises that run AI implementation as a procurement exercise buy the tool but not the workflow redesign, training, and governance the tool needs — and the contract misses the clauses that matter: notification before material model updates, retention and exit-deletion terms, indemnity for AI outputs, audit rights over the vendor's own AI integrations. Two incidents show what those missing clauses cost: a vendor silently added generative AI to the US National Eating Disorders Association's approved rule-based helpline bot, which then gave harmful advice, and the Salesloft Drift breach spread through a vendor's own integrations. Clearing this blocker means rewriting the procurement pathway for AI purchases, because pushing more of them through the old pathway reproduces the same gaps.

Future Processing, 2025; CNN on NEDA Tessa, 1 Jun 2023; blueprint Part IV on contract controls.

Blocker: the unbuilt labour agreement

In co-determination jurisdictions, the workforce path runs through negotiated agreements. German works councils hold a mandatory co-determination right — not merely consultation — over technical systems capable of monitoring employee behaviour or performance, which covers most AI tools in human-resources contexts; the employer must reach agreement before introduction. A rollout plan that discovers this after deployment has found its wall. In Singapore the same institutional layer is an asset rather than an obstacle: Company Training Committees under the National Trades Union Congress give firms a ready-made joint forum where AI adoption, job redesign, and training are negotiated together.

Clearing the hinge: identify, own, act

The Performance Bridge, Chapter 12

Identify. List every blocker, then find the one that controls the others: not what makes execution harder, but what makes it impossible or close to it. If you removed this one, would the rest become manageable? That is the critical blocker. Until it is removed, improving everything else produces zero improvement in outcomes.

Own. The highest authority in the business owns blocker removal — the person who can reallocate budget, change reporting lines, override resistance, and make decisions that stick. The transformation lead identifies and escalates; the senior leader removes. If that leader treats blockers as someone else's problem, the bridge does not become operational.

Act. Set a deadline and treat it as non-negotiable. Blockers that are acknowledged but not acted on become permanent: people learn to work around them, workarounds become culture, and the blocker calcifies into "the way things are done here". Once the critical blocker is cleared, a new constraint will surface; that is the sequence operating as designed, because removing one constraint exposes the next one.

Execution in reality

With the hinge cleared, the bottom-up flow can carry the design: structure gets built and used, redesigned workflows run, drivers start moving, and outcomes begin to materialise against their baselines. The evidence signature of a cleared hinge is simple — the same drivers that sat flat during the blocked period begin to respond to actions. If they still do not, either another wall is standing or the driver set itself failed its tests.

Before any of this works, three prerequisites must hold (The Performance Bridge, Chapter 22). Applied in the wrong environment, the model does not fail gracefully — it makes things worse: measurement becomes surveillance, drivers become quotas, and gaming multiplies. The three checks below form a gate. Passing it is Phase 0 work; failing it means fixing the foundation before building on it.

Readiness check Five culture questions Data maturity Levels 0 to 3 Sponsor commitment Asked and recorded in advance The gate Any check failed? Fix that first Proceed to the roadmap All three checks pass

Select a check to open its questions; the teal box jumps to the roadmap.

The readiness check

The Performance Bridge, Chapter 22

The bridge rests on a foundation of empowerment; in command-and-control environments, drivers become surveillance and teams learn to hit metrics rather than improve outcomes. An AI programme on the same foundation produces mandates and performative usage. Five questions, translated for AI:

  1. Can teams adjust how they use AI without seeking approval for every variation?
  2. When an AI metric moves the wrong way, do leaders ask what happened, or who is to blame?
  3. Do teams see their own usage and quality data in real time, or is it locked in leadership reports?
  4. Is there room to experiment with different tools and workflows within guardrails?
  5. When someone surfaces an AI failure — a hallucinated answer sent to a customer, a near-miss data leak — are they thanked or punished?

The fifth question carries extra weight in Asia-Pacific deployments, where face-saving and hierarchy dynamics demonstrably suppress error reporting: an incident process that depends on frontline staff volunteering bad news will underreport unless leaders de-personalise error and visibly reward the report. If readiness is missing, build safety and decision space first — the methodology can wait.

The data maturity check

The Performance Bridge, Chapter 22

Driver management depends on data: a driver can only be watched if it is measured, and only measured if the process captures the information in the first place. The four maturity levels below govern two distinct AI questions: whether the AI itself can be grounded in reliable data, and — more often missed — whether the outcome can ever be validated. A pilot in a process with no baseline data cannot be proved to have worked or failed, which is precisely the condition behind the many pilots claiming success nobody could demonstrate.

LevelStateImplication for the AI transition
0 — No dataThe information is not capturedThe first AI investment is instrumenting the process; alternatively find a measurable proxy and label it as one
1 — Raw dataExists somewhere; manual extractionAcceptable for a pilot, but cannot be kept up at scale — automate before scaling
2 — Accessible dataAvailable in reports, with lagWorkable; do not let perfect data block good-enough validation
3 — Dashboard dataReal time, self-serviceIdeal for driver management; rare for new processes

The trap: teams select ideal drivers they cannot track, then quietly substitute weaker ones because the data is available. Never accept a weak driver merely because it is easy to measure — and remember Gartner's prediction that organisations will abandon 60 per cent of AI projects unsupported by AI-ready data.

The sponsor question

The Performance Bridge, Chapters 22 and 23

Every sponsor says they are committed; commitment reveals itself only under pressure — when the transformation conflicts with something the sponsor values more, or when holding an influential person accountable would cost a relationship. Ask directly before starting, and record the answer: when this transformation creates conflict with other priorities, will you personally engage to resolve it? When obstacles require your authority to remove, will you remove them?

The book's claim that weak sponsorship dooms well-designed transformations now has quantitative support: where the chief executive is personally accountable for decisions based on AI outputs, confidence in AI strategy runs at 60 per cent versus 22 per cent, meaningful business value at 57 per cent versus 21 per cent, and established returns at 14 per cent versus 4 per cent.

When a check fails

A failed check is useful information: it says what must be fixed before the transition can work, and proceeding anyway is how the model makes things worse. The remedies map to the failed check:

  • Readiness failed: build psychological safety and decision space first. Model curiosity about misses, democratise access to performance data, and delegate method decisions within a defined scope. Sophisticated measurement of an unaddressed dysfunction produces dashboards and gaming, not outcomes.
  • Data maturity failed: either invest in data capture as a structural requirement before the transformation, run a deliberately small pilot on manual extraction while automating, or accept a documented proxy temporarily. Whichever path, capture the baseline before deploying anything.
  • Sponsorship failed: do not compensate with methodology — no framework substitutes for a sponsor unwilling to spend political capital. Either renegotiate the mandate or shrink the scope to what the available authority can actually clear.

The roadmap is the Performance Bridge built top-down and run bottom-up, phased. Durations are indicative for a mid-sized or large organisation; for a small or medium enterprise the same sequence compresses from quarters into weeks — select the final row for that variant.

Phase 0 · foundations First weeks Diagnose and protect Strategy Test, readiness, Tier 1 controls Phase 1 · design First quarter Baseline and drivers Lighthouse domain, baseline, charter Phase 2 · pilot Months 3 to 9 Pilot to production Assist-first pilots, workflow redesign Phase 3 · scale Month 9 onward Scale through structure Portfolio, governance rhythm, watch list The SME variant The same sequence, compressed into weeks

Select a phase to open its full checklist.

Phase 0 — foundations

First weeks, before any pilot
  1. Run the Strategy Test. Write the diagnosis: what specific obstacle is AI meant to overcome, in which one to three business domains? If no obstacle can be named, stop — any spend at this point is fear-of-missing-out spend.
  2. Run the readiness and data-maturity checks. If blame culture would poison incident reporting, build safety first. Map candidate domains against the data-maturity levels; where a process has no baseline data, instrumenting it is the first AI investment, because a pilot without a baseline is unfalsifiable.
  3. Ask the sponsor question directly and record the answer: when this transformation creates conflict with other priorities, will you personally engage to resolve it? The KPMG accountability data shows why this is not a formality — chief-executive accountability roughly triples the rate of realised value.
  4. Deploy Tier 1 controls immediately: a written acceptable-use policy naming permitted tools and prohibited data classes; an approved-tool list with a sanctioned, convenient default; business-tier subscriptions with no-training defaults; a staff briefing built on the two incidents that make the risks concrete (Samsung engineers pasting source code into a consumer chatbot, and the deepfake video call that extracted US$25.6 million from the engineering firm Arup); and out-of-band verification with dual authorisation for payment instructions. These cost little and close the largest immediate exposures — and they honour the finding that bans fail while sanctioned channels work.
Blueprint Parts III and IV; KPMG, Jun 2026; Dezeen on Arup, 17 May 2024.

Phase 1 — diagnose, baseline, and design

Roughly the first quarter
  1. Select one to three lighthouse domains — the small number of business areas where the approach will be proved end to end — prioritising workflows that are high-volume, well-bounded, and instrumented, the task shape where the experimental literature shows AI reliably helps.
  2. Capture roughly 90 days of baseline data on the exact metrics the AI is meant to move.
  3. Build the driver set before the pilot: two to four drivers per initiative, each passing the four tests, including at least one counter-driver and one vitality driver, with commitment levels assigned. Run the Gaming Check on every driver: how might someone hit this number without improving the outcome?
  4. Write the pilot charter: success criteria as measurable business metrics against the baseline; kill criteria set in advance ("if the AI absorbs less than 15 per cent of the baselined work by day 80, or exception rates exceed the current error rate, we shut it down"); adversarial testing included; the human escalation path designed as a feature; one named owner.
  5. Clear the hinge: list the blockers, identify the critical one, and have the senior leader remove it on a deadline before execution begins.
Blueprint Parts II, III, and VI; kill-criteria example per Summit Trails.

Phase 2 — pilot to production

Roughly months three to nine
  1. Run production-intent pilots with real users in a controlled slice, at the assist-first autonomy level, long enough to validate quality before scaling — DBS ran nine months; the NHS ran a nine-site, 17,000-encounter trial before committing to a regional rollout.
  2. Redesign the workflow around what the pilot shows. This is the Value Action — the book's term for the one action that directly delivers the value and cannot be skipped — and the strongest predictor of eventual financial impact; fund the people-and-process work at something approaching the 70 per cent of effort the evidence says it needs.
  3. Review drivers weekly, outcomes at the gate. Where the counter-driver diverges — cost falling while quality falls — treat it as gaming or quality erosion and investigate before scaling. Kill weak pilots on the pre-agreed criteria and say so plainly; a disciplined exit is a success of method.
  4. Stand up Tier 2 controls and the governance rhythm: the use-case registry with risk tiering, vendor assessments with model-change and retention clauses, human sign-off for high-stakes outputs, and the cross-functional committee cadence.
  5. Channel bottom-up energy onto the sanctioned platform: open safe tooling to employees, celebrate the internal builders, and let the shadow economy surface into governance.
Blueprint Parts II, IV, and VIII; GOSH nine-site trial.

Phase 3 — scale and sustain

Months nine onward
  1. Scale what validated, through structure. Convert pilot practices into embedded processes: a standing dashboard in place of periodic reminders, a defined role in place of one heroic champion, a curriculum in place of one-off workshops. Test every element by asking what would keep working if the AI champion left tomorrow.
  2. Manage AI as a portfolio rather than as single projects: a stage-gated pipeline mixing quick wins, process transformations, and a few strategic bets; a single benefits register per workforce population to prevent double-counting; honest reporting of the adoption ramp.
  3. Report outcomes so they would survive adversarial review: against the public baseline and target, validated with a counterfactual, alongside the vitality drivers. Announce nothing as achieved that is only projected.
  4. Keep the Watch List on the leadership agenda — the book's list of conditions no structure can install: trust, fear, psychological safety around error reporting, the judgment to override the AI, and alertness to the moment when a control itself has become the obstacle. These require sustained leadership attention, because they erode when attention moves elsewhere.
  5. Track the regulatory trajectory in every operating jurisdiction, aligning with voluntary frameworks before they harden — the consistent pattern across Singapore, South Korea, Australia, and the European Union, and the cheapest compliance strategy on record.
Blueprint Parts IV, V, and VIII; Biswas, The Performance Bridge (2026), Ch. 13 on stickiness.

The SME variant — the same sequence in weeks

For small and medium enterprises

Every link stays, but each is kept short. The owner writes the one-paragraph diagnosis of the costliest bottleneck. A funded trial — in Singapore, through the GenAI Sandbox or the Productivity Solutions Grant — tests one pre-approved tool against one workflow. The driver set is one operational number plus one counter-driver. Tier 1 controls are the whole security programme: business-tier subscriptions, a one-page acceptable-use policy, and the payment-verification procedure. The continuation decision follows the measured result inside the trial window.

The evidence that the compressed discipline works: Singapore's first sandbox cohort enrolled more than 150 SMEs and approximately 80 per cent continued using their solution after the funded period ended, and the worked example — Far East Flora's pre-approved support chatbot — reported a 67 per cent reduction in staff hours spent answering queries. Among Singapore's AI-using firms, 84 per cent rely on off-the-shelf tools: the SME on-ramp is buy rather than build, one workflow at a time, with one honest metric.

Security and intellectual-property controls are Structure: they make safe actions possible at scale. The diagnosis behind this whole tab is that ungoverned adoption converts AI's productivity upside into leakage, fraud, and legal exposure — and bans demonstrably fail, so the problem to solve is a governed path as convenient as the shadow path. The tiers build on each other; each assumes the one before it is in place.

The threat landscape, 2023–2026 Shadow AI Prompt injection Agentic compromise Deepfake fraud IP risk, both directions Tier 1 · foundational Every company, in weeks Policy and a sanctioned path Acceptable use, payment verification Tier 2 · managed Once AI use is broad Contracts, inventory, sign-off Data-loss prevention, vendor checks Tier 3 · advanced Regulated and agentic Isolation and red-teaming Least privilege, evals, certification The lethal trifecta and the Agents Rule of Two The agentic threat model, and the design gate that bounds it The framework stack NIST, ISO 42001, OWASP, EU AI Act, Singapore

Select a threat to see the incident evidence, a tier to see its control list, or the bottom row for the frameworks to align with.

Threat: shadow AI

Employees using unapproved AI tools is the single most consistent empirical finding of the period, and the behavioural baseline every control programme has to assume. More than 80 per cent of employees report using unauthorised AI tools, and 68 per cent of security leaders admit doing so themselves. The financial consequence is quantified: breaches involving shadow AI cost an average of US$4.63 million — about US$670,000 more than breaches without it — and among organisations suffering AI-related incidents, 97 per cent lacked proper AI access controls and 63 per cent had no AI governance policy.

Samsung (2023) — the formative case

Within weeks of the electronics group permitting the consumer chatbot ChatGPT, engineers pasted semiconductor source code and meeting records into it in three separate incidents. Every incident was a well-intentioned attempt to work faster rather than an act of theft — which is exactly why prohibition alone fails: the motivation it targets is not the motivation driving the behaviour.

The design conclusion: shadow AI is evidence that employees perceive real value and the organisation failed to provide a sanctioned path. Roughly 35 per cent of workers say they would keep using unauthorised tools under an explicit ban. Every control in the tiers below is therefore paired with a legitimate, convenient alternative.

Threat: prompt injection

Content that manipulates a model into ignoring its instructions is ranked the number-one risk in the OWASP (Open Web Application Security Project) Top 10 for large-language-model applications, and it is structural to current models — a standing property, not an edge case.

EchoLeak (June 2025) — the escalation point

The first documented zero-click prompt-injection exploit against a production AI system: a single crafted email to a Microsoft 365 Copilot user could trigger remote exfiltration of tenant data when the user later asked a routine question. Patched server-side, but the attack class remains.

The customer-facing versions

A car dealership's chatbot was manipulated through its prompt into "agreeing" to sell a US$76,000 vehicle for one dollar, and the parcel carrier DPD's chatbot was induced to swear at a customer after an update shipped without adversarial regression tests. Assume every customer-facing or email-reading AI system is adversarially probed from day one, and give no bot any authority — over pricing, commitments, or refunds — that the business is not prepared to honour.

The lethal trifecta, and the Agents Rule of Two

The agentic threat model, and the design gate that bounds it

Security researcher Simon Willison named the combination that makes an AI agent exploitable: access to private data, exposure to untrusted content, and the ability to communicate externally. An agent holding all three can be steered by a poisoned input into exfiltrating whatever it can read — no software vulnerability required, because the vulnerability is the design itself. EchoLeak is the trifecta operating exactly as described: Copilot could read tenant data, ingested untrusted email, and could emit content externally, and a single crafted message connected the three.

Prompt injection remains an unsolved problem — models cannot yet reliably distinguish instructions from data — so the defence is architectural rather than instructional: remove one of the three properties from any agent that does not need all of them, instead of writing ever-sterner system prompts. The Replit incident, in which an autonomous coding agent deleted a production database despite direct instructions not to touch it, is the same lesson in miniature: the instruction failed to constrain the agent, and only removing the write credential would have.

The Rule of Two as a checkable gate

Meta's security team operationalised the trifecta as a design rule: within a session, an agent should satisfy no more than two of three properties — it can process untrustworthy inputs, it can access sensitive systems or private data, it can change state or communicate externally. If a task genuinely requires all three without a fresh session, the agent should not operate autonomously; it requires supervision through human-in-the-loop approval or another reliable validation. The rule becomes a one-line question at use-case registration: which two of the three properties does this agent hold, and who approved the third?

Property heldExampleWhat to remove or gate
Untrusted inputs + private dataAn assistant summarising inbound email against the customer recordNo external send or state change without human approval
Untrusted inputs + external actionsA public-facing chatbot that can raise ticketsNo access to sensitive systems; no authority the business will not honour
Private data + external actionsAn internal agent drafting and sending reports from company dataNo untrusted content in its context — curated sources only

Two caveats belong in any briefing. Meta positions the rule as a supplement to, not a substitute for, least privilege and defence in depth — designs satisfying it can still fail against other threat vectors. And the rule bounds the blast radius of prompt injection rather than preventing the injection itself, which is precisely why it is honest: it accepts that the attack cannot yet be stopped and removes what the attack could reach. In Performance Bridge terms it is a pure structure control — a permission boundary that persists without ongoing attention, where a system-prompt instruction is a reminder that fails silently.

Threat: agentic compromise

The shift from chatbots to agents moved the attack surface from answers to actions. Every AI integration holding standing credentials into core systems is a new trust relationship that inherits none of the perimeter's controls.

Salesloft Drift (August 2025)

Attackers compromised Drift, an AI chat agent widely integrated with the Salesforce platform, stole the standing access tokens it used to connect to customers' systems, and read data across more than 700 organisations without touching any victim's own network.

AI-orchestrated espionage (November 2025)

Anthropic reported disrupting a state-sponsored campaign in which an agentic coding tool was manipulated into executing an estimated 80–90 per cent of an intrusion campaign against roughly 30 targets. Attacker cost per intrusion is falling faster than defender cost.

Controls: least-privilege, short-lived credentials for every agent integration; environment separation and human approval gates for destructive actions, because the Replit incident showed that instructions alone do not constrain an agent; and Singapore's Cyber Security Agency has published a dedicated addendum on securing agentic AI.

Threat: deepfake fraud

Arup, Hong Kong (January 2024) — the anchor case

A finance employee at the British engineering firm Arup joined a video conference in which the chief financial officer and several recognisable colleagues appeared and spoke — every participant except the victim was an AI deepfake built from public footage. Reassured, the employee executed fifteen transfers totalling HK$200 million (about US$25.6 million). The fraud surfaced only when the employee followed up with the real head office.

The control that failed was a payment process, and the fix is also a process. The payment-authorisation procedure accepted a video call as identity verification. Out-of-band callback verification and dual authorisation for large transfers — ordinary process controls that involve no AI technology — would have defeated the fraud. Voice cloning now requires only seconds of audio, which every executive who has spoken at a conference or on an earnings call has already supplied. This belongs in Tier 1 for every company regardless of size.

Threat: intellectual-property risk, in both directions

Direction one: company IP leaking into third-party models

The commercial tier and the signed contract are the control. Business and enterprise tiers of the major providers do not train on customer content by default; consumer tiers commonly do. Two caveats belong in every policy briefing: litigation can override retention promises — a court order compelled production of 20 million anonymised consumer conversation logs in the New York Times case — and the protections are contractual, not architectural: the same model reached through a personal account carries none of them.

Direction two: the IP status of AI outputs

The US Copyright Office confirmed that human authorship is required and that purely AI-generated output is not copyrightable — so fully AI-generated marketing assets or content libraries may be unprotectable against copying by competitors. On training data, the industry pattern is settlement and licensing rather than definitive rulings (the authors' class action against Anthropic settled for US$1.5 billion), which leaves residual uncertainty with user-side companies and argues for vendors offering indemnification. For AI-generated code, indemnity is conditional — Microsoft's commitment requires duplicate-detection filters kept enabled — and suggestions match training-set code in roughly 1 per cent of cases, a real surface for open-source licence contamination at enterprise scale.

Tier 1 — foundational

Every company, including small firms, within weeks
  • A written acceptable-use policy naming permitted tools and prohibited data classes.
  • An approved-tool list with a sanctioned, convenient default — so employees have a legitimate path as easy as the shadow one.
  • Business-tier subscriptions with no-training defaults; no confidential data in consumer tools.
  • A staff briefing built on the two incidents that make the risks concrete: the Samsung engineers who pasted source code into a consumer chatbot, and the Arup finance employee who transferred US$25.6 million after a deepfake video call.
  • Out-of-band verification and dual authorisation for payment instructions, regardless of how convincing the requester appears.

Risks addressed: shadow AI, casual data leakage, deepfake payment fraud. For a small or medium enterprise, this tier plus vendor attestations is the entire security programme — and it addresses the highest-loss incident types of the period.

Tier 2 — managed

Once AI use is broad; assumes Tier 1
  • Enterprise agreements with negotiated retention, subprocessor, and exit-deletion clauses.
  • Data-classification gates specifying which classes may enter which tools, enforced by data-loss-prevention tooling tuned for AI destinations.
  • An AI inventory and registration process covering vendor-embedded AI — the estate you did not know you had.
  • Vendor risk assessment extending to the vendor's own AI integrations, because the Salesloft Drift breach spread through exactly that route — with model-change notification clauses, because the National Eating Disorders Association's vendor changed an approved bot's behaviour without telling it.
  • Human sign-off for high-stakes outputs — public content, legal, financial, clinical — reflecting the Air Canada liability principle that the company owns what its AI says.

Risks addressed: contractual exposure, an unknown AI estate, third-party AI supply chain, output liability.

Tier 3 — advanced

Regulated sectors, high-value IP, agentic deployments; assumes Tiers 1 and 2
  • Private endpoints or virtual-private-cloud deployment for sensitive workloads; AI gateways as a policy-enforcing control plane.
  • Systematic red-teaming and evaluations against OWASP and MITRE ATLAS scenarios before and after every deployment and update.
  • An AI-specific incident-response playbook: model rollback, guardrail-regression detection, prompt-injection triage.
  • Least-privilege, short-lived credentials for every agent integration; environment separation; approval gates for destructive actions.
  • The Agents Rule of Two applied as a design gate at use-case registration: no agent session holds untrusted inputs, sensitive data access, and external actions at once without human-in-the-loop supervision (the lethal trifecta panel).
  • NIST AI Risk Management Framework alignment internally, and ISO/IEC 42001 certification where customers demand attestation.

Risks addressed: prompt injection, agentic compromise, regulatory attestation. Two design principles govern all three tiers: every control is paired with a sanctioned path, because prohibition alone demonstrably fails; and the highest-loss incidents of the period (Arup, Salesloft Drift) were or would have been defeated by process controls, not AI-specific technology.

The framework stack

Four voluntary frameworks plus two binding regimes cover most of what a transitioning company needs to align with. The common pattern is NIST for internal discipline and ISO for external attestation.

FrameworkWhat it gives you
NIST AI Risk Management Framework + Generative AI Profile (July 2024)Four functions — govern, map, measure, manage — twelve generative-AI risk categories, and more than 200 mapped actions; the inventory requirement underpins registry practice
ISO/IEC 42001:2023The certifiable AI management-system standard, now a procurement gate — AWS, Anthropic, Snowflake, Salesforce, and ServiceNow certified in sequence, and it increasingly appears in due-diligence questionnaires
OWASP Top 10 for LLM Applications and MITRE ATLASThe practitioner threat taxonomy and the red-teaming scenario source
EU AI Act (in force 1 August 2024, staged to 2027–2028)Extraterritorial: a Singapore firm whose AI-assisted hiring or credit decisions touch EU individuals is in scope without any EU presence, with penalties up to EUR 35 million or 7 per cent of global turnover
Singapore's stack: PDPC guidelines, IMDA Model AI Governance Framework for Generative AI, CSA secure-by-design guidelinesVoluntary but functioning as a maturity ladder — and the regional direction is voluntary frameworks hardening into obligations, so early alignment is the cheapest compliance strategy

Sector overlays: in healthcare, every ambient-documentation vendor is a business associate requiring signed agreements, consent is now in litigation, and an AI that informs clinical decisions is likely regulated as a medical device; Singapore's refreshed AI in Healthcare Guidelines (AIHGle 2.0, March 2026) map developer, deployer, and user responsibilities. In financial services, the Monetary Authority of Singapore has moved from voluntary principles to a consultation on binding AI risk-management guidelines for all financial institutions.

Roughly 70 per cent of the effort in a successful AI transformation belongs in people and process change, yet change management is the part of most programmes that is communicated rather than measured. On the Performance Bridge, change management is the vitality dimension of the transition — the trust, skills, energy, and willingness to report problems that let the organisation produce the same results again — and it is run with the same driver discipline as everything else. Select any element to open the evidence behind it.

Performance drivers What the period achieves Vitality drivers What keeps results repeatable Every performance driver gets a vitality counter-driver — select either box The vitality driver catalogue for an AI transition Trust in AI outputs Only 46 per cent are willing to trust Displacement fear Highest where adoption is highest Shadow-usage rate A verdict on the sanctioned offer Incident reporting rate An empty register is the warning Human-override rate Read in both directions Rework burden on recipients Workslop and code churn Redeployment of freed time Where do the saved hours go? Skill progression Beyond the first plateau How to read vitality drivers Three interpretive rules that prevent the classic misreadings The vitality bankruptcy pattern Performance bought with an unwatched vitality debit What working looks like The four moves in the success cases The one-page plan Five checkable lines

Select a driver for its measurement, its evidence base, and the documented failure it would have caught early.

Two kinds of outcome, one pairing rule

The frame: change management is the vitality dimension, measured
The Performance Bridge, Chapters 1, 10, and 23

The book distinguishes performance outcomes — the results of the period — from vitality outcomes: the organisation's sustained capability to keep producing results, carried in trust, skills, energy, and the willingness to surface problems. A transition that hits its performance targets while burning vitality has borrowed its results rather than earned them, and "burning vitality for performance" is one of the six mistakes that break the bridge.

The operating rule that makes this measurable is pairing: every performance driver in an AI programme gets a vitality counter-driver that makes its gaming, or its collateral damage, visible. The payments company Klarna illustrates the stakes: its cost per contact fell while customer satisfaction and escalation behaviour deteriorated, and because only cost was on the dashboard, the divergence stayed invisible for a year until it forced a public reversal. A paired dashboard would have shown it within weeks.

Performance driverVitality counter-driver
Average handle time, resolution speedFirst-contact resolution; re-contact rate within 24 hours
Cost per contactCustomer satisfaction on AI-touched journeys; complaint and escalation rates
AI-assisted output volumeRework rate and downstream recipient ratings; code-churn ratio
Adoption rate under a mandateOutput-quality audits on sampled work; trust and sentiment pulse
Sources: Biswas, The Performance Bridge (2026), Ch. 1, 10, 23; Forbes on the Klarna reversal, 18 May 2025.

Trust in AI outputs

What it detects: whether people rely on the tools or quietly route around them

The global baseline is poor, which is why a programme that never measures trust is assuming a condition the evidence says is absent: a 48,000-respondent study across 47 countries found 66 per cent of people using AI regularly but only 46 per cent willing to trust it, and 48 per cent of employees admitting to using AI in ways that contravene company policy.

How to measure it. A monthly pulse with specific behavioural questions — "Would you send an AI draft to a client without rewriting it?" — rather than an annual engagement survey. The annual instrument fails the book's WATCH test: it arrives months after the damage, far too late to steer.

Displacement fear

What it detects: whether the workforce believes the honest version of the programme's intent

The Asia-Pacific paradox makes this the region's central change-management driver: 53 per cent of frontline workers fear job loss — against 36 per cent globally — in the same survey that shows the region leading the world in adoption, with Singapore, South Korea, and Thailand reporting the highest concern. In Singapore, only 15 per cent of employees feel confident about job security.

How to measure it. The same anonymous pulse, plus attrition and internal-transfer patterns in the teams the programme touches. What answers the fear is a demonstrated redeployment path rather than reassurance: IKEA retrained 8,500 call-centre workers into a measured design-advisory channel, and the Italian bank Intesa Sanpaolo negotiated its workforce transition with its unions. A town-hall promise is only a projection, and the workforce knows it.

Commonwealth Bank of Australia — the lasting cost of a false claim

The bank announced redundancies on an unvalidated projection about its voice bot and reversed them within weeks when the union disproved the claim — and the union observed that the damage to remaining staff's trust survived the reversal.

Shadow-usage rate

What it detects: whether the sanctioned path is good enough

More than 80 per cent of employees report using unapproved AI tools at work, 68 per cent of security leaders admit doing so themselves, and roughly 35 per cent of workers say they would keep using unauthorised tools under an explicit ban. Read correctly, a high shadow-usage rate tells you the sanctioned tools are less useful or less convenient than the unsanctioned ones — and it also measures a workforce appetite that the successful programmes put to work rather than suppressed.

How to measure it. The share of observed AI traffic flowing through sanctioned tools — a structural metric — plus a periodic anonymous self-report. The successful cases show the driver moving when the sanctioned offer improves: employees of the Spanish bank BBVA built more than 2,900 custom assistants within months of receiving a governed platform, and staff at DBS, Singapore's largest bank, have built roughly 26,000 personal agents on its internal platform.

Incident and near-miss reporting rate

What it detects: whether psychological safety exists — the condition all AI governance depends on

Every incident-response plan, model registry, and governance committee depends on one behaviour: a person voluntarily reporting that the AI got something wrong. No procedure can produce that behaviour; it appears only where people believe reporting is safe. Practitioners across Asia consistently report that employees rarely volunteer bad news unless leaders make it explicitly safe, and healthcare research links inclusive leadership, through psychological safety, to error reporting.

How to measure it. The count of voluntarily reported AI errors and near-misses, read against usage volume — and read with the inversion rule: early in a transition, a rising count usually means the reporting channel works. The alarming signal is an empty incident register beside heavy usage, because that combination measures suppressed reporting rather than safe operation. De-personalise error — attribute the failure to the system rather than to the person who surfaced it — and visibly reward the report.

Sources: HRM Asia on psychological safety in Asia; PMC, inclusive leadership and error reporting, 2022; Biswas, The Performance Bridge (2026), Ch. 14 and 22.

Human-override rate

What it detects: whether autonomy matches demonstrated reliability — in both directions

The share of AI outputs overridden or heavily edited by the responsible human is one of the few drivers that must be read two-sided. Persistently high override rates mean the tool is not trusted or not good enough for the workflow it sits in. Near-zero override on consequential decisions signals automation complacency — the condition that let a chatbot's invented bereavement policy reach an Air Canada customer, and an invented licensing policy reach Cursor's users, unchecked.

The third reading is saturation: when alert volume is high enough, override becomes reflexive dismissal. The Epic sepsis model — a widely deployed hospital early-warning system — showed this at scale: external validation found it missed most sepsis cases while alerting so constantly that clinicians learned to dismiss it. The driver only works when someone owns the question of which of the three readings applies.

Sources: Moffatt v. Air Canada, 2024 BCCRT 149; The Register on the Cursor support bot; Biswas, The Performance Bridge (2026), Ch. 8.

Rework burden on recipients

What it detects: whether AI is transferring effort rather than removing it

"Workslop" — AI-generated content that masquerades as good work but lacks the substance to advance a task — is the named version of this driver's failure state: 40 per cent of surveyed workers had received it in the previous month, each incident cost recipients nearly two hours of rework, the estimated invisible tax runs to roughly US$186 per employee per month, and recipients marked the sender down as less capable and less trustworthy. The survey was self-selected, so treat the numbers as suggestive; the concept is the useful part.

The code version is quantified more cleanly: duplicated code blocks rose eightfold in 2024 and code churn — lines reverted or substantially revised within two weeks — rose from 3.1 to 5.7 per cent of changed lines. More AI-assisted output shipped; more of it came straight back.

How to measure it. Sampled downstream ratings of AI-assisted work and rework-time tracking, paired as the counter-driver to any output-volume metric.

Redeployment of freed time

What it detects: whether "hours saved" is becoming value or evaporating

Time saved is a legitimate leading indicator, but it is a proxy — value materialises only when the freed time is deliberately redeployed, and the book's surrogation warning says teams forget proxies are proxies and optimise the number itself. The evidence says the forgetting is the norm: in a 3,200-respondent study, companies were more likely to reinvest AI time savings in more technology (39 per cent) than in employee development (30 per cent), 32 per cent simply increased workload, and nearly four in every ten gained hours were lost to reworking AI output.

How to measure it. Track where the freed hours verifiably went — development, redesigned work, or simply more volume. An organisation reporting "hours saved" as value, without a redeployment mechanism and a destination measurement, is optimising a proxy that has decoupled from the outcome.

Sources: Workday, 14 January 2026; Biswas, The Performance Bridge (2026), Ch. 9.

Skill progression

What it detects: whether capability is compounding or plateauing at novice prompting

MIT's GenAI Divide report — the study behind the widely quoted finding that 95 per cent of generative-AI pilots showed no measurable profit impact — diagnosed the cause as a learning gap: organisations stall because neither their tools nor their people advance beyond the first plateau. Training hours completed will not catch this, because hours accumulate without learning. The driver that does is the number of trained employees observed applying the skill in their own workflow, tracked over time, with internal certification progression as the long-run signal.

DBS — training built as standing structure

The Singapore bank runs a mandatory generative-AI curriculum for all employees with role-specific depth, and had built years of data-literacy programmes ("Data Heroes") before the generative wave — so the skill baseline the transition needed already existed when the technology arrived.

Sources: MIT Project NANDA, "The GenAI Divide", July 2025; DBS annual report and newsroom, cited in the case library.

How to read vitality drivers

Three interpretive rules, briefed before the first dashboard review

Rule one: early in a transition, rising incident reports are good news. A climbing count in the first months usually means the reporting channel works and people believe using it is safe. The untrained reading — "incidents are up, the programme is in trouble" — punishes exactly the behaviour the programme needs, and in face-conscious cultures it will silence reporting permanently. State the inversion explicitly to every leader who will see the dashboard.

Rule two: vitality drivers are the counter-drivers of the performance set. Pairing each performance driver with a vitality driver is what makes the Klarna pattern — cost falling while satisfaction, escalations, and complaints deteriorate — detectable by design: a paired dashboard shows the divergence in weeks, where a cost-only dashboard concealed it until the company had to reverse in public.

Rule three: treat silence as a measurement gap, never as consent. A workforce can be anxious, adopting, and silent all at once — 58 per cent of Asia-Pacific respondents would use AI without company approval while half fear for their jobs. Where hierarchy or face dynamics suppress individual voice, make the pulse survey anonymous and the questions specific, and use the institutional channels — Singapore's Company Training Committees, European works councils, unions — as the collective route for what individuals will not say alone.

Sources: Biswas, The Performance Bridge (2026), Ch. 10 and 14; BCG, 30 October 2025.

The vitality bankruptcy pattern

The failure shape: a performance result purchased by an unwatched vitality debit
The Performance Bridge, Chapter 23 — Mistake 4, burning vitality for performance

The clearest change-management failures of 2024–2026 share one shape: the performance number was real and loudly announced, the vitality account drained unmeasured, and the debt was called in publicly.

Klarna — the clearest case

The Swedish payments company replaced the work of roughly 700 customer-service agents with an AI assistant whose cost performance was real; the quality of resolution, and customers' inability to reach a human, were the account nobody watched. The chief executive's reversal statement — that letting cost dominate the evaluation had produced lower-quality service — is a description of this mistake by a leader who had just made it.

Commonwealth Bank of Australia — the workforce version

The bank cut roles on an unvalidated projection about its voice bot and reversed the cuts within weeks when the union disproved the claim. The union noted that the damage to remaining staff's trust survived the reversal: the correction did not restore what the false claim had spent.

Taco Bell — the customer-side signal at its most literal

Customers of the fast-food chain deliberately ordered 18,000 cups of water from its drive-thru voice AI to force a human onto the ordering line. When customers sabotage the system to reach a person, the trust measure has gone from negative to adversarial, and the escalation path has failed as a piece of design.

Mandates and bans are the same mistake in opposite directions. Usage mandates — Shopify made AI use a performance-review criterion, and Meta announced "AI-driven impact" as a core performance expectation — buy visible adoption at the price of performative usage where training and redesign are absent, because once usage is graded, the graded quantity is what employees inflate. Bans buy visible compliance at the price of hidden usage: Samsung banned generative AI after its source-code leaks, and roughly 35 per cent of workers say they would defy an explicit ban. Both approaches substitute an instruction for the work of enablement and trust, and both optimise a compliance metric while the underlying vitality drains.

What working change management looks like

The four moves the success cases share — each on a bridge link

Sanctioned capability before any demand (Structure). Panasonic Connect built a secure internal assistant that removed the compliance ambiguity around usage before asking anything of employees, paired it with visible top-management commitment, and measured honestly: 788,000 hours saved in fiscal 2025 against a stated long-term goal. DBS built its internal tools and training curriculum on the same logic — the safe path existed before usage was expected.

Channel the energy that already exists (Actions). The shadow-AI statistics, read correctly, measure appetite. The 2,900 assistants built by BBVA's employees and the 26,000 personal agents built by DBS staff did not come from persuasion campaigns; both companies provided a governed outlet, and demand that already existed flowed into it.

Answer displacement fear with a demonstrated path (Outcomes). IKEA's reskilled call-centre workers and the union-negotiated transition at the Italian bank Intesa Sanpaolo could point to named colleagues whose work changed and survived — which is evidence in the Outcomes sense, where a promise is only a projection. The message must also fit the labour context: in Japan, where the workforce problem is shortage, the credible framing is that AI covers the work there are no longer people for; in Singapore, where displacement fear is among the world's highest, the message must address that fear directly.

Leaders model what no policy can install (the Watch List — the book's list of conditions structure cannot solve). DBS describes the hardest part of its transformation as cultural: moving a banking culture built on flawless execution to one that experiments, framed internally as winning "heads and hearts". The leadership behaviours are the lever for every driver structure cannot move: visibly using the tools, receiving bad news about the AI well, and thanking the person who reports the near-miss.

Sources: Panasonic Newsroom; OpenAI on BBVA; PYMNTS on IKEA; MIT Sloan Management Review on DBS; Biswas, The Performance Bridge (2026), Ch. 14.

The change-management plan as a bridge artefact

One page, five lines, every line checkable
  1. A named owner and a budget line. The 70 per cent of effort that belongs in people and process gets an owner and funded time; a plan without one restates the inverse allocation — 80 per cent of budget on platforms — as intent.
  2. Enablement sequenced before any demand. Sanctioned tool, training with protected time to learn, and at least one redesigned workflow exist before usage appears in anyone's performance expectations.
  3. Two or three vitality drivers with commitment levels, paired against the initiative's performance drivers and reviewed in the same governance rhythm as the security controls, so the operating review cannot ignore them. Using the book's three commitment levels — GUARDRAIL for what must never be breached, COMMITTED for what the team is expected to hit, STRETCH for genuine ambitions — incident-reporting suppression sits at GUARDRAIL, trust and redeployment targets at COMMITTED, and cultural transformation of the DBS depth is honestly a STRETCH.
  4. The interpretive rules briefed to every dashboard reader — in particular the inversion on incident reporting — before the first review meeting rather than after the first misreading.
  5. A demonstrated redeployment path for the roles the programme touches, stated in Outcomes terms: named roles, destination work, and a validation check.
Sources: Biswas, The Performance Bridge (2026), Ch. 11, 13, 14; BCG, the 10-20-70 allocation.

The business case is where the Outcomes and Strategy links meet the chief financial officer. The 2024–2026 shift is from experimentation budgets to profit-and-loss accountability: only 7 per cent of leaders have established returns, payback expectations run two to four years, and leaders with strong cost visibility are five times more likely to achieve ROI (return on investment). A case that survives scrutiny contains eight components — select each for its evidence.

1 · Diagnosis and counterfactual 2 · A measured baseline 3 · Total cost of ownership 4 · The benefit ledger, named 5 · Scenario discipline 6 · An adoption ramp 7 · Kill criteria and stage gates 8 · A named accountable owner The three time-value ledgers Utilisation, capacity, headcount — different claims Two worked examples Service deflection, ambient clinical documentation The six failure patterns How AI business cases inflate

Select a component, a ledger, an example, or the failure patterns.

1 · A strategic diagnosis with a quantified counterfactual

The case opens with the business problem — the same "what's in the way?" that opens the Strategy link — and a quantified do-nothing scenario. Without a counterfactual, the funding request is a technology purchase; with one, it is a comparison a chief financial officer can actually make. The failure statistics of the period describe programmes funded without this component: pilots whose success could never be demonstrated because nobody stated what failure would have looked like.

2 · A measured baseline

Roughly 90 days of pre-deployment data on the exact metrics the AI is meant to move. A pilot in a process with no baseline is unfalsifiable — precisely the condition under which the MIT GenAI Divide report found pilots claiming success nobody could demonstrate. The gold standard is an experimental design: the economists Brynjolfsson, Li, and Raymond used a staggered rollout across 5,179 support agents as a natural control group, measured issues resolved per hour — a business metric rather than a self-report — and found productivity up 14 per cent on average and 34 per cent for novice agents. Even a simple matched comparison group is far better than no counterfactual at all.

3 · A full total-cost-of-ownership model

Gartner has estimated real generative-AI deployment costs at US$5 million to US$20 million, and 96 per cent of deploying organisations faced higher-than-expected costs. The model must itemise: licences and seats; token and inference consumption at projected volumes (agentic workflows consume 5 to 30 times more tokens per task than a chatbot query — the cost line is a function of usage, never a fixed fee); integration engineering; data preparation; security and compliance review; monitoring and evaluation infrastructure; training; and sustained change management. The hidden 70 per cent is the people-and-process line the 10-20-70 rule demands — and vendor-embedded AI has produced cost uplifts of around 30 per cent on existing tools without upfront disclosure. On sourcing: 76 per cent of enterprises now buy rather than build, and bought solutions convert pilot to production at 47 per cent against 25 per cent; the working rule is to buy the platform, build only the differentiating layer, and treat integration as the main engineering effort.

4 · Benefits stated in one named ledger

The single most common inflation in AI business cases is conflating the three time-value ledgers — utilisation, capacity, and headcount avoidance (see the teal row below for the full treatment). The case must name which claim it is making and the conversion mechanism: freed time redeployed to what, hiring avoided against which demand forecast, or which named workforce plan. Quality effects go on both sides: repeat-inquiry reductions are a credit; rework of AI output (nearly two hours per workslop incident; repeat contacts at double cost behind inflated deflection rates) is a debit most cases omit.

5 · Scenario discipline

A conservative case at partial benefit realisation, with assumptions documented per scenario. The empirical justification is the skew of AI returns: Deloitte's 2024 wave found about 20 per cent of organisations achieving returns above 30 per cent on their most advanced initiative while the median struggled — a distribution that punishes single-point estimates. The skew is also the argument for funding a portfolio (quick wins, process transformations, a few strategic bets) rather than one flagship: a single bet maximises variance, not expected value.

6 · An adoption ramp

A case that books full benefit in the first quarter is contradicted by the entire survey base: 91 per cent of finance organisations report only low-to-moderate initial impact, with impact compounding over years of adoption maturity — organisations further along the curve are two to three times more likely to report moderate or high impact. Most executives expect two to four years to satisfactory returns; only 6 per cent achieve payback in under a year. Model the ramp, and count the early capability building — data pipelines, evaluation infrastructure, AI-literate staff — as option value that lowers the cost of every subsequent use case.

7 · Pre-agreed success metrics, kill criteria, and stage gates

Success criteria defined before the pilot begins, as measurable business metrics against the baseline — "reduce average handle time from 8.2 minutes to 5.5 minutes", not a satisfaction score. Kill criteria set in advance are what make the funding request credible: "if the AI absorbs less than 15 per cent of the baselined work by day 80, or exception rates exceed the current error rate, we shut it down." A disciplined exit is a success of method — McDonald's two-year, hundred-restaurant drive-thru test ended cleanly when accuracy never reached the threshold at which removing the human paid. Stage gates run quick wins, then pilots with go and no-go decisions, then scale, with weak pilots actively retired at each gate.

8 · A named accountable owner

The empirical justification for refusing to fund orphan initiatives: in organisations where the chief executive is personally accountable for decisions based on AI outputs, confidence in AI strategy runs at 60 per cent versus 22 per cent elsewhere, meaningful business value at 57 per cent versus 21 per cent, and established returns at 14 per cent versus 4 per cent. Leaders with strong cost visibility are five times more likely to achieve returns (15 per cent versus 3 per cent) — which also argues for a single AI investment register, since 42 per cent of leaders report only partial visibility of AI spending and the same productivity gain is routinely double-counted across initiatives.

The three time-value ledgers

Time saved multiplied by loaded labour cost is the dominant benefit claim in AI business cases, and it needs three separate ledgers that must not be conflated, because they are different claims with different evidence requirements.

LedgerThe claimValued only if
(a) UtilisationThe same people do the same work fasterThe freed time is redeployed to specified activities — which does not happen by default: companies are more likely to reinvest savings in more technology (39 per cent) than in people (30 per cent), 32 per cent simply increase workload, and nearly four of every ten saved hours are lost to reworking AI output
(b) CapacityThe team absorbs volume growth without hiringHiring avoided is valued against a demand forecast. Klarna's famous "work of 700 agents" was this claim — workload equivalence during a growth phase, not dismissals
(c) Headcount reductionPayroll actually fallsA named workforce plan exists — and the validation burden is the one the Commonwealth Bank reversal established: an adversarial party will examine the evidence

The wider calibration: workers using AI save on average around 5 per cent of weekly hours; self-reports overstate the savings (in one randomised trial, developers believed they were 20 per cent faster while measuring 19 per cent slower); and a Danish study matching 25,000 workers to administrative records found time savings of about 3 per cent of hours with no measurable effect on earnings. Task-level gains become firm-level gains only when someone deliberately redesigns what the freed time is spent on — which is what Verizon did by redirecting it into retention and sales conversations.

Two worked examples

Service operations: customer-service deflection economics

The unit economics: human-handled tickets average US$8–12 (US$25–35 in business-to-business software support) against roughly US$0.50–1.05 for AI-handled tickets, with realistic containment of 40–60 per cent for well-documented tier-one query types. The critical measurement distinction is deflection against containment: deflection counts sessions that ended in the bot; containment counts contacts fully resolved with no follow-up on any channel within 24 hours. True containment typically runs 15–25 percentage points below reported deflection, and the customers who call back cost roughly double. The credible case prices benefits on containment, nets off repeat contacts, and measures by category before scaling — as the Singapore telecommunications company Singtel did, committing to scale only after measuring about 73 per cent containment on troubleshooting queries and 76 per cent on roaming sign-ups.

Healthcare: ambient clinical documentation

The evidence is deliberately two-sided and the honest case reflects it. Time savings are real but modest per encounter: Kaiser Permanente's 15,791 hours across roughly 2.5 million encounters, and 16 minutes of documentation per eight hours of care in the five-centre academic study. Burnout effects are the strongest measured outcomes — a 21.2 per cent reduction in burnout prevalence at Mass General Brigham. But the Peterson Health Technology Institute's multi-system evaluation concluded that clear financial returns have not yet been demonstrated, with licence costs of roughly US$100–600 per provider per month. The honest business case in this domain is therefore a retention-and-capacity case with option value — clinician time, burnout, and turnover-cost avoidance modelled explicitly as assumptions — and not a payroll-savings case.

The six failure patterns

  1. Benefits claimed on gross time savings. Multiplying minutes saved by salary assumes 100 per cent of freed time converts to value; the telemetry and survey evidence shows it is reinvested diffusely, and self-reports overstate the savings to begin with.
  2. The adoption ramp ignored. Full benefit booked in quarter one against a survey base showing years-long compounding.
  3. Rework and verification costs omitted. Nearly two hours per workslop incident; repeat contacts at double cost behind inflated deflection rates.
  4. Double-counting across initiatives. The same "10 per cent productivity gain" for the same population claimed by multiple projects; the control is a single benefits register per workforce population.
  5. No counterfactual. Pre-and-post comparisons without a control group attribute secular trends — volume changes, seasonality, other process changes — to the AI.
  6. Benefits priced on vendor-defined metrics. Where the vendor bills per "resolution" or "outcome", the buyer must independently audit that billable outcomes match business outcomes — the deflection-versus-containment problem applies directly to invoices.
Blueprint Part VI; outcome-pricing context per Intercom Fin pricing and Salesforce Agentforce repricing, 15 May 2025.

A multinational cannot run one global AI rollout playbook: labour institutions, regulatory regimes, and cultural dynamics change the transition path region by region. Select a region for its profile, then browse the case library below — every case is drawn from the blueprint's verified research, colour-coded by what it teaches.

Singapore and APAC Europe and the UK North America Middle East Latin America What travels

Select a region, or the teal box for the cross-regional lessons.

Singapore and Asia-Pacific

The most institutionally supported environment

Singapore is the case where the Structure link is substantially built by the state: the National AI Strategy 2.0 (December 2023) with more than S$1 billion committed at Budget 2024; the IMDA (Infocomm Media Development Authority) Model AI Governance Framework for Generative AI with its nine dimensions; the AI Verify Foundation's testing tools including Project Moonshot; sector guidance from MAS (Monetary Authority of Singapore) for finance — on a visible trajectory from voluntary principles to binding risk-management guidelines — and AIHGle 2.0 for healthcare; and an SME grant architecture (Productivity Solutions Grant, GenAI Sandbox, Digital Leaders) whose first sandbox cohort saw roughly 80 per cent of participating SMEs continue after funding ended. The caveat: subsidy has not closed the SME divide (14.5 per cent against 62.5 per cent large-firm adoption), because the binding constraint is absorptive capacity, not cost.

The central change-management fact: the adoption-anxiety paradox

APAC leads the world simultaneously in adoption and in fear: 78 per cent of employees use AI at least weekly (versus 72 per cent globally), yet 53 per cent of frontline workers fear job loss (versus 36 per cent), with Singapore, South Korea, and Thailand highest — and 58 per cent would use AI without company approval. Only 15 per cent of Singapore employees feel confident about job security. The change-management task this creates is unusual: the enthusiasm already exists, so the work is converting existing, anxious, partly ungoverned energy into governed adoption.

Culture-aware design rules

  • Senior sponsorship carries more weight where hierarchy is accepted — and so does its liability: pair endorsement with an explicit, repeated invitation of dissent.
  • Face-saving suppresses the error reporting AI governance depends on; de-personalise error and reward the report.
  • Kiasu — Singapore's fear of losing out — plausibly drives both fast adoption and reluctance to be the visible first failure; treat it as an interpretive lens and make early pilots low-stakes.
  • Institutional voice channels compensate for suppressed individual voice: NTUC's Company Training Committees tie AI adoption to negotiated wage and career outcomes, converting individual anxiety into collective voice.
  • Japan is the mirror image — lowest adoption (51 per cent) but lowest job fear (40 per cent) — because labour institutions make displacement unlikely; the credible message there is that AI covers work there are no longer people for, while in Singapore the message must address displacement directly.
DBS Bank — culture before models

The hardest part on the record was not the technology but moving a flawless-execution banking culture to one that experiments — framed internally as winning "heads and hearts". Panasonic Connect shows the same levers working in Japan: a sanctioned secure tool, visible top-management commitment, and measured outcomes (788,000 hours saved in fiscal 2025).

Europe and the United Kingdom

The adoption story is better than its reputation

On official all-firm surveys, EU enterprise adoption (20.0 per cent in 2025, up 6.5 points in a single year — the fastest rise recorded) matches or exceeds the comparable US Census figure, and Denmark (42 per cent) exceeds anything measured in the US on the same instrument. Where Europe genuinely lags is capital deployed and the share of firms scaling beyond pilots: 48 per cent of Europe's largest companies had scaled a transformational generative-AI initiative against 31 per cent of smaller ones — and Europe's economy is small-firm-heavy.

What a deployer must actually do

The EU AI Act's near-term duties for a typical deploying enterprise are AI literacy, prohibited-use screening, and supplier due diligence; the heavier high-risk conformity work (employment screening, credit scoring) arrives on the staged timeline. The UK remains principles-based through sector regulators, with legislation repeatedly deferred. The deeper structural difference is labour institutions: German works councils hold mandatory co-determination over AI systems capable of monitoring performance, so rollouts are negotiated, slower, and redeployment-heavy.

The redeployment-first pattern

IKEA's operating company retrained 8,500 call-centre workers as remote design advisers feeding a €1.3 billion sales channel (with the caveat that the channel's revenue is not causally attributable to the retrained cohort alone), and the Italian bank Intesa Sanpaolo negotiated 9,000 exits alongside 3,500 technology hires through its unions, attributing an incremental €100 million of 2025 gross income to its AI programme. In Europe's negotiated-labour environment, a redeployment plan is the institutional price of adoption, agreed before rollout rather than offered afterwards.

The NHS — evaluate, then procure

The UK's National Health Service ran a nine-site London trial of an ambient clinical scribe across more than 17,000 patient encounters — measuring 23.5 per cent more direct patient-interaction time and 13.4 per cent more emergency-department patients seen per shift — before committing to a phased rollout toward 20,000 clinicians and a national supplier registry. It is the public-sector procurement model worth copying: evidence first, contract second.

North America

The capital lead, and the regulatory whipsaw

US private AI investment runs roughly 23 times China's, and US generative-AI investment exceeds China and Europe combined — but on official all-firm surveys US adoption (17–20 per cent) no longer leads the EU. The federal direction reversed twice in two years: a January 2025 executive order revoked the prior administration's AI order, and a December 2025 order created a Department of Justice task force to sue states over their AI laws, while Colorado gutted and delayed its landmark state act in May 2026. The operative constraint for a deployer is litigation risk and a volatile state patchwork — employment discrimination (the iTutorGroup settlement; the Workday collective action reaching the vendor itself), consumer protection, and product liability — rather than a single rule.

The reference deployments, and the claim inflation

The strongest governed rollouts are here — JPMorgan's single secure gateway for 200,000 users, Morgan Stanley's evaluation-first assistant, Kaiser Permanente's published ambient-scribe study, Cleveland Clinic's year-long vendor bake-off — alongside the sharpest natural experiment: Providence deployed comparable scribe technology and reached only about 8 per cent active clinician use. This is also the home of the unaudited chief-executive productivity claim: Amazon's "4,500 developer-years saved" and Salesforce's "AI does 30 to 50 per cent of the work" should be discounted against third-party-evaluated results; Accenture's US$5.9 billion in generative-AI bookings against US$2.7 billion recognised revenue is the cleaner demand indicator — client commitments run ahead of delivered work.

Middle East

The Gulf pattern is state-led and infrastructure-first, with the state simultaneously investor, customer, and regulator. Saudi Arabia's HUMAIN carries a reported US$77 billion infrastructure plan targeting 6.6 gigawatts of data-centre capacity by 2034; the UAE's Stargate project targets 1 gigawatt; and a November 2025 US export authorisation for up to 70,000 advanced chips unblocked frozen capital. Regulation runs through procurement and certification rather than statute: Dubai's AI Seal — a six-tier certification with 325 corporate applicants by mid-2025 — is positioned to become the credential for bidding on government AI work, and the UAE Cabinet has announced that half of government services will move to autonomous AI systems within two years. Corporate deployments follow the sovereignty logic: Aramco built its industrial large language model in-house on decades of proprietary drilling data, and ADNOC signed a US$340 million agentic-AI contract for its upstream value chain. For a vendor or partner, certification readiness is a market-entry requirement, not an afterthought.

Latin America

The region shows high grassroots usage against a severe investment deficit: 14 per cent of global visits to AI tools with only 11 per cent of global internet users, yet just 1.12 per cent of global AI investment against 6.6 per cent of global output. Brazil's PL 2338 — an EU-style risk-based bill with fines up to R$50 million and a contested training-data remuneration requirement — passed the Senate in December 2024 and remains before the Chamber of Deputies. The flagship corporate deployments run on imported foundation models but are genuinely instructive: Nubank cut chat response time 70 per cent with roughly 55 per cent of tier-one enquiries resolved by its assistant across more than 2 million monthly chats, and Mercado Libre built an internal AI platform ("Verdi") that lets its 17,000 developers compose AI workflows — a platform-layer answer to the sprawl of unconnected AI agents.

What travels, and what is context-dependent

Four patterns recur in every region

  1. Staged rollouts with measured gates beat big-bang deployments. BBVA's ramp — 3,300 licences, then 11,000, then all 120,000 employees, each stage gated on measured usage — is the counter-model to Klarna-style substitution: adoption first, redesign second.
  2. The rollout design determines adoption, because the technology is broadly comparable. The US health systems Kaiser Permanente and Providence bought similar ambient-scribe products; one reached 7,260 actively using physicians, the other roughly 8 per cent of its clinicians.
  3. Two use cases have replicated, quantified outcomes across regions: ambient clinical documentation and first-line customer service — everywhere from California to London to São Paulo.
  4. Vendor and chief-executive claims systematically outrun independent evaluation. Weight peer-reviewed and third-party-evaluated results; treat the rest as marketing until validated.

The multinational minimum

An EU AI Act compliance baseline applied group-wide (the strictest broadly applicable regime often becomes the de facto global standard); country-level workforce processes wherever works councils or sector unions exist; US deployments reviewed for state-law exposure and litigation risk; and certification and evaluation-registry readiness for government business in the Gulf or the UK. For Singapore and APAC companies expanding outward, the asymmetry runs in their favour on speed but against them on compliance depth — Europe's staged obligations and negotiated-labour environment are the two systems most unlike their home experience.

The cross-case analysis: where the chain breaks

What the 43 documented cases show, read together on the five links

Every case in the library below carries a five-segment verdict strip: for each link of the Performance Bridge, the public record either shows the case ran that link well, shows the chain broke there, or offers no evidence either way. Counting those verdicts across all 43 cases produces the table below, and the heat map above the library shows the same tallies as bars. Three findings stand out.

LinkRan wellChain broke hereNot evidenced in the public record
Structure
Actions
Drivers
Outcomes
Strategy

First, success is a complete chain, not a strong link. Fourteen of the twenty-one validated successes ran all five links well — the winning profile is not excellence at one link but the absence of a broken one. The successes share the same construction: structure built before demands were made, workflows redesigned around the tool, a small honest driver set, and outcomes validated against a baseline someone else could check.

Second, failures compound. Thirteen of the failure and caution cases broke at two or more links, because a break upstream travels: a missing diagnosis leaves nothing to validate, and missing structure turns ordinary employee enthusiasm into leakage. The links that broke most often are Structure and Drivers — eleven cases each — which are exactly the two links the book identifies as the seat of leadership leverage, and the two that most programmes neglect.

Third, the most common condition at Outcomes is silence. Outcomes shows fewer outright breaks than Structure or Drivers, but it is the link most often marked "not evidenced": in nineteen of forty-three cases, the public record contains no validated outcome at all — no baseline, no audited result, no independent check. An organisation that cannot show its outcome has not necessarily failed, but it cannot demonstrate success either, and the aggregate statistics of the period suggest most of that silence is not modesty.

The verdicts are editorial judgments applied to the published record, case by case; the reasoning for each is in the case's expandable reading below, and the "Diagnose it yourself" mode hides these verdicts so you can form your own before comparing.

The case library

Border colour: validated success failure or reversal caution — claims unaudited, outcomes mixed, or in flight
Spine strip on each card — the five links in causal order (S structure, A actions, D drivers, O outcomes, St strategy): the case ran this link well the chain broke here not evidenced either way. Coloured segments open the matching panel.
Region
Lesson
Sector
Broke at
DBS Bankbanking · Singapore

The anchor case: a public S$1 billion economic-value target set in 2022 and met in financial year 2025, on a governed data platform, the PURE governance gate, and roughly 26,000 employee-built agents.

DBS customer-service assistantcontact centre · Singapore

A nine-month pilot before scaling: near-100 per cent transcription accuracy, up to 20 per cent call-time reduction, close to 90 per cent positive officer feedback. The assistant supports the human officer inside the call rather than replacing the officer.

Singtel and Sierratelecommunications · Singapore

The telecommunications company measured containment by query category before scaling: about 73 per cent of troubleshooting queries and 76 per cent of roaming sign-ups completed without a human.

Commonwealth Bank scam programmebanking · Australia

Customer scam losses fell 76 per cent from their peak, independently reported — a single external harm metric, pointed at the moment of intervention.

Commonwealth Bank redundancy reversalbanking · Australia

Forty-five roles were cut on projected call-volume reductions; the union showed volumes had risen, and the bank reversed the cuts and apologised. The projection had been treated as if it were a validated result.

Deloitte Australia refundprofessional services · Australia

AI-fabricated citations in a government assurance report, found by an external researcher; part of the fee refunded. Citation verification is cheap; its absence voided the deliverable.

Samsung data leaksemiconductors · South Korea

Engineers pasted semiconductor source code into a consumer chatbot within weeks of it being permitted — well-intentioned shadow use filling a structural vacuum.

Arup deepfake fraudengineering · Hong Kong

A finance employee at the engineering firm transferred HK$200 million after a video call in which every other participant was a deepfake. The control that failed was the payment process itself: no out-of-band verification step existed.

NUHS and SingHealthhealthcare · Singapore

Singapore's public healthcare clusters — the National University Health System (NUHS) and SingHealth — run a sovereign, self-hosted clinical model (RUSSELL-GPT) and ambient documentation on the national Synapxe health-technology platform, so the data-sovereignty blocker was removed once, nationally, for every cluster.

Panasonic Connectelectronics · Japan

A sanctioned internal assistant for all 12,500 Japanese employees, deployed explicitly to reduce shadow-AI risk; 788,000 hours saved in fiscal 2025 — 3.4 per cent of working hours.

Far East Floraretail SME · Singapore

The Singapore florist deployed a pre-approved chatbot under the IMDA sandbox and measured a 67 per cent reduction in staff hours spent on customer queries. It is the template for a small firm: one tool, one workflow, one honest metric.

OCBC and UOBbanking · Singapore

An instructive contrast in defensible defaults: OCBC built its own governed chatbot and scaled to all 30,000 staff fast; UOB piloted a bought tool with 300 users first. Company-reported gains; treat magnitudes with care.

Klarnafintech · Sweden

The payments company's 2024 containment numbers were real but validated on cost and volume only; quality declined on complex cases and the company publicly reversed into the hybrid design it could have started with. It has become the reference case the whole field calibrates against.

BBVAbanking · Spain

The staged ramp: 3,300 licences, then 11,000, then all 120,000 employees, each stage gated on measured usage — adoption first, redesign second.

IKEA / Ingkaretail · Netherlands and Sweden

8,500 call-centre workers were retrained as remote design advisers feeding a €1.3 billion sales channel, with the caveat that the channel's revenue is not attributable to the retrained cohort alone. In Europe's labour environment, a redeployment plan of this kind is the negotiated price of automation.

Intesa Sanpaolobanking · Italy

A union-negotiated generational transition — 9,000 exits, 3,500 technology hires — with an incremental €100 million of 2025 gross income attributed to 117 live AI applications.

NHS ambient-scribe trialhealthcare · United Kingdom

Led by Great Ormond Street Hospital, the National Health Service trialled an ambient scribe across nine London sites and 17,000 patient encounters (measuring 23.5 per cent more patient-interaction time) before committing to a phased 20,000-clinician rollout and a national supplier registry — evaluation first, procurement second.

Octopus Energy / Krakenutilities · United Kingdom

The British energy retailer climbed the autonomy ladder in order: assisted drafting proved quality against the human baseline before any autonomous answering, and the autonomous stage was still measured head-to-head against human responses.

INGbanking · Netherlands

A quantified bottleneck (16,500 customers queuing weekly), a seven-week guarded build, daily regression testing on 500 real chats, and 20 per cent more customers helped without escalation.

BMW Plant Regensburgmanufacturing · Germany

Predictive maintenance trained on the plant's own fault data, validated on a metric the plant already tracked: around 500 minutes of avoided line disruption per year.

DPD chatbotlogistics · United Kingdom

A guardrail regression shipped in an update with no adversarial regression test; the bot swore at a customer and the story went global. Release discipline applies to bots.

Babylon Healthhealthcare · United Kingdom

A US$4.2 billion valuation on clinical claims that outran the evidence; bankruptcy in 2023 left health systems that had built pathways on it holding the risk. Vendor viability is part of clinical risk.

Lufthansaaviation · Germany

4,000 mostly administrative roles to go by 2030, explicitly coupled to AI and negotiated through employee representatives — announced structure, outcomes still prospective.

JPMorgan LLM Suitebanking · United States

One governed large-language-model (LLM) gateway for the whole bank — about 200,000 users onboarded in eight months — built deliberately to prevent a thousand ungoverned pilots from starting instead.

Morgan Stanleywealth management · United States

The wealth-management firm graded the assistant's answers against expert answers before rollout, over a curated research library. The result was over 98 per cent adoption among advisor teams, with document access rising from roughly 20 to 80 per cent of the library.

Verizontelecommunications · United States

An assistant for 28,000 service representatives was tuned until it could answer 95 per cent of questions, and the freed time was deliberately redirected into retention and sales conversations — a commercial return on top of the cost saving.

Kaiser Permanentehealthcare · United States

The largest published ambient-scribe deployment: 7,260 physicians, 2.5 million encounters, results in NEJM Catalyst — voluntary adoption, quality assurance, peer-reviewed validation.

Cleveland Clinichealthcare · United States

A year-long pilot across 80-plus specialties and 25,000 encounters before choosing a vendor: 32 per cent more patient face time, 49.6 per cent less after-hours documentation.

Providencehealthcare · United States

The US health system deployed scribe technology comparable to Kaiser Permanente's and reached roughly 8 per cent active clinician use. The two systems form a natural experiment showing that the rollout design, rather than the tool, determined adoption.

Air Canada chatbotaviation · Canada

The liability precedent: a tribunal rejected the claim that the chatbot was "a separate legal entity" — the company owns what its AI says.

McDonald's–IBM drive-thrufood service · United States

A bounded two-year test of automated drive-thru ordering across a hundred restaurants, measured against operational reality and exited cleanly when accuracy never reached the level at which removing the human paid. The disciplined exit is itself the lesson: the method succeeded even though the technology did not.

Taco Bell voice AIfood service · United States

Customers ordered 18,000 cups of water from the drive-thru voice AI to force a human onto the line. The escalation path had failed as a piece of design, and the deployment's purpose was inverted.

Replit agentsoftware · United States

An autonomous coding agent on the Replit platform deleted a production database against direct instructions, then produced output that masked the damage. Only a permission boundary enforced in the system — never an instruction in a prompt — reliably constrains an agent.

Cursor support botsoftware · United States

An ungrounded support bot invented a policy and triggered cancellations — the Air Canada failure repeated by an AI-literate company a year after the ruling.

NYC MyCity chatbotgovernment · United States

Illegal advice to small businesses, defended with a disclaimer. In regulated domains a wrong answer is a compliance event, and a disclaimer does not transfer the risk.

Epic sepsis modelhealthcare · United States

Deployed on vendor performance claims; external validation found it missed roughly two-thirds of cases while alerting constantly. Vendor numbers are marketing until validated locally.

NEDA Tessahealthcare · United States

The US National Eating Disorders Association (NEDA) replaced its helpline with a rule-based bot, Tessa; the vendor silently added generative AI, and the bot gave harmful dieting advice to vulnerable callers. Unannounced vendor model changes are a first-order governance risk.

iTutorGroup and Mobley v. Workdayhiring · United States

The first AI hiring-discrimination settlement, and a collective action reaching the screening vendor itself — liability attaches to the algorithm's operators, not only its buyers.

Amazon Q and Salesforce Agentforce claimstechnology · United States

"4,500 developer-years saved"; "AI does 30 to 50 per cent of the work" — unaudited executive claims that shifted under scrutiny. Discount vendor-side numbers; weight third-party evaluation.

Nubankbanking · Brazil

Chat response time down 70 per cent, roughly 55 per cent of tier-one enquiries resolved across 2 million-plus monthly chats, with human escalation preserved.

Mercado Libre Verdie-commerce · Argentina and Brazil

An internal AI platform for 17,000 developers — making AI a capability any team can assemble, the platform-layer answer to agent sprawl.

Aramco and ADNOCenergy · Saudi Arabia and UAE

Sovereignty-first industrial AI: a proprietary model built on decades of drilling data, and a US$340 million agentic contract for the upstream value chain — at-scale bets whose outcomes are still being proven.

UAE government servicespublic sector · UAE

A target of half of government services on autonomous AI within two years, steered through procurement certification — the most aggressive public-sector adoption programme anywhere, in flight.

Glossary of abbreviations
AI
Artificial intelligence
APAC
Asia-Pacific
CSA
Cyber Security Agency of Singapore
EBIT
Earnings before interest and taxes
EHR
Electronic health record
EU AI Act
The European Union's Artificial Intelligence Act (in force August 2024, staged application)
FM1–FM8
The eight recurring failure modes identified in the 2023–2026 case record; the full catalogue is the bottom row of the Failure modes tab
FOMO
Fear of missing out
GPAI
General-purpose AI (a category under the EU AI Act)
IMDA
Infocomm Media Development Authority (Singapore)
IP
Intellectual property
ISO/IEC 42001
The certifiable international standard for AI management systems
LLM
Large language model
MAS
Monetary Authority of Singapore
NHS
National Health Service (United Kingdom)
NIST
US National Institute of Standards and Technology (publisher of the AI Risk Management Framework)
NTUC
National Trades Union Congress (Singapore)
OWASP
Open Web Application Security Project
PDPC
Personal Data Protection Commission (Singapore)
ROI
Return on investment
SME
Small and medium enterprise
TCO
Total cost of ownership