A widely-cited McKinsey number puts roughly 80% of AI initiatives as failing to reach production. The number is approximately right, and what it under-reports is harder: of the 20% that do ship, a meaningful share are quietly switched off within twelve months, replaced by the previous workflow, or kept running purely so nobody has to admit they didn't work.
This essay is for the executive who has either lived through that or is about to. It is not about which model to pick. It is about the five conditions that determine, before any engineering work begins, whether the pilot has any chance of mattering. Each of these patterns shows up over and over in the work, across MediaTech, FinTech, HealthTech, RetailTech, and the rest. Each of them is fixable. Most of them are fixable before a single line of code is written.
The single most useful framing I have for this work: most failed AI pilots are not technology failures. They are organisational failures that the technology revealed.
— Pattern 01The outcome isn't defined sharply enough to fail.
The first conversation usually starts something like this: "We want to do something with AI." Or: "The board has asked us to develop an AI strategy." Or, slightly more advanced: "We want to use AI to drive efficiency."
None of these are outcomes. They are anxieties dressed as initiatives. The question that exposes the problem is simple: "What is the one number you would expect to move six months after we start, and what is that number today?" A surprising fraction of leadership teams cannot answer it without a long internal conversation. Some cannot answer it at all.
This matters because an outcome that cannot be stated cannot fail. A pilot pointed at "drive efficiency" produces a deck twelve months later showing some efficiency was, in fact, driven somewhere. Nobody can say it failed. Nobody can say it succeeded. The pilot continues consuming budget because there is no clean way to kill it.
What to do instead. Before any technical scoping, write the constraint document: one page that names the single business outcome, the measurable baseline today, the target six and twelve months out, and the go/no-go criterion at the end of the first phase. The exercise itself is most of the value. Half the time, the act of trying to write it surfaces that the leadership team does not actually agree on what they are trying to change. That disagreement is the first thing to fix, and no amount of model selection will fix it.
"Drive efficiency" fails. "Reduce customer onboarding time from 47 days median to under 14 days, without increasing customer support headcount, by Q4" passes. The difference is not pedantry — it is whether the project has a way to end.
The data isn't where the model needs it to be.
Every AI engagement uncovers a data layer that is more fragmented than the leadership team believed. This is a near-universal pattern. The CTO will, with complete sincerity, describe a unified data warehouse. The data engineers will, with equal sincerity, describe seven systems with three competing definitions of "customer."
Both are right. The warehouse exists; it just doesn't contain the data the AI use case actually needs in the form it needs it. The customer records are unified for analytics but not for real-time decisioning. The labeled data sits in a CSV on one analyst's laptop. The compliance team's classification taxonomy is six months out of sync with product's.
The pilot starts anyway. Three months in, 60% of the engineering team's time is being spent on data plumbing rather than the model. The plumbing work is unglamorous, invisible to the executive sponsor, and produces no tangible output. The pilot enters the dangerous phase where the team is working hard, the sponsor is anxious, and nothing is visibly progressing.
What to do instead. Audit the data before the model selection conversation. The audit is small — usually two weeks — and produces three artifacts: an inventory of the data the use case actually requires, a quality and labeling assessment of each source, and an honest estimate of cleanup work as a percentage of total engagement effort. Plan around that estimate, not in spite of it. The teams that get this right routinely find that the right first project is not the headline AI initiative — it is a focused data cleanup that makes the AI initiative feasible six months later.
Half the time, writing the constraint document surfaces that the leadership team does not agree on what they're trying to change.
The org structure isn't set up to absorb the output.
This is the failure mode that catches the most sophisticated organisations. The technology works. The data is fine. The model ships. And then, six weeks after launch, adoption is sitting at 12% and nobody can explain why.
The reason is almost always one of three structural conditions. Either there is no named executive owner whose review or bonus depends on the outcome — the project is owned by a committee, which means it is owned by nobody. Or the team building the AI sits organisationally distant from the team using it, and a trust gap has formed before launch that the model itself cannot close. Or — most common — the people who would use the AI output have never been involved in shaping how it presents itself, and the interaction pattern feels imposed rather than designed.
In healthcare engagements this presents very clearly: a model that is statistically better than human readers produces zero adoption because the interface asks clinicians to defer, and senior clinicians have spent decades correctly training themselves never to defer to a tool they cannot reason about. The fix is not a better model. The fix is changing the interaction shape so the clinician retains authority and the AI surfaces context. It is a UI decision masquerading as an adoption problem.
What to do instead. Before the build phase begins, name the owner. One person. Their performance review must be tied to the outcome — not the launch, the outcome. Then bring the eventual users into the Design phase as co-designers, not as user-research subjects. The disagreement loop with users during design is what produces both the right interaction shape and the trust that lets adoption happen.
— Pattern 04The pilot has no go/no-go criteria.
The fourth pattern is the financial one. A pilot starts with vaguely defined success and even more vaguely defined failure. Six months in, results are mixed. The team has built real capability, learned real things, and not yet produced the headline outcome. The question is whether to continue.
Without explicit go/no-go criteria written down at the start, this conversation almost always resolves in favour of continuation. The reasons are predictable. The team has invested significantly; sunk-cost dynamics are real. The executive sponsor doesn't want to call a failure on their own initiative. The vendor or consultant has incentive to keep the engagement going. The narrative shifts from "did we hit the target?" to "look at all the things we learned."
The pilot becomes a permanent line item. Two years later it is still consuming budget, no longer aspiring to its original outcome, kept alive by the absence of a clean exit. This is how AI initiatives die — not by failing visibly, but by living quietly forever.
What to do instead. Write the go/no-go criterion before the work begins, sign it, and structure the engagement so it has gates. Three to four weeks in: did Diagnose produce a sharper constraint than we had at the start? If yes, continue. If no, stop. Eight to ten weeks in: did Design produce a buildable solution the eventual owner is willing to commit to operating? If yes, continue. If no, stop. The phases are not just delivery milestones — they are off-ramps. The team that designed them must also have the authority to pull them.
— Pattern 05The model is treated as the work, instead of 20% of it.
The fifth pattern is subtler than the others, and the one most likely to catch teams that have done one or two AI initiatives before. The team plans the engagement around the model build. They budget for the model build. They communicate progress in model performance metrics. They celebrate when the model hits target accuracy.
And then the project enters production, and the model is roughly 20% of what determines whether it lands. The deployment pipeline matters. The monitoring layer matters. The drift detection matters. The integration with downstream systems matters. The training program for operators matters. The runbook for the operations team matters. The escalation paths for when the model is uncertain matter. The post-launch UX feedback loop matters.
Most of this is invisible until you need it, and then it is the difference between a model that quietly degrades into uselessness over six months and a capability that compounds. The teams that ship and then keep shipping have built this layer deliberately. The teams that ship once and then watch their initiative fade have not.
What to do instead. Plan the engagement in four phases — Diagnose, Design, Build, Embed — not in a single phase called "the project." The Embed phase is where the model becomes a capability the organisation actually owns. It is also the phase most likely to be cut for time, which is exactly why it fails so often. Build embed in from the start; price it distinctly; assign it to someone whose job is not "ship the model" but "make the model land."
So what does this add up to?
Two things, mostly. First: the technology questions in AI work — which model, which framework, which infrastructure — are now the easy part. They are commodified, well-documented, and rarely the difference between success and failure. The hard part is everything around the technology: the outcome definition, the data foundation, the organisational ownership, the disciplined gating, and the embed work. These are not technical skills. They are operating skills.
Second: the firms and teams that consistently land AI work are the ones who have learned to front-load the hard part. Two to four weeks of Diagnose before a single line of code is written. A constraint document signed by the leadership team before scope is finalised. An owner named before kick-off. Go/no-go criteria written into the contract. An Embed phase priced and staffed as carefully as the Build phase.
This is unglamorous. It is also the difference between an AI initiative that earns its budget and one that quietly burns it. The work is the work — and most of it happens before the model does.