Strategy
Why most AI pilots produce nothing
The failure is almost never the technology. It is buying access before building judgment, choosing the visible use case over the valuable one, and running a pilot that was never defined tightly enough to succeed or fail.
The short answer
MIT reviewed over 300 disclosed AI initiatives and found roughly 95% delivered zero measurable return. The differentiator was approach, not model quality or regulation.
The most common pattern we see is sequence: organizations buy multiple platform subscriptions before investing in any literacy, then attribute the resulting poor outcomes to the technology.
Budgets concentrate in sales and marketing because those use cases are easy to imagine and pitch, while clearer returns tend to appear in unglamorous back-office work.
A pilot that can succeed has one named business outcome, a baseline measured beforehand, an owner who does the work, and a defined stopping point.
MIT's Media Lab reviewed more than 300 publicly disclosed AI initiatives, conducted 52 organizational interviews, and surveyed 153 executives. The headline finding was that approximately 95% of these initiatives showed zero measurable return, and that the difference between the 5% and everyone else was not explained by model quality or regulatory environment. It came down to approach.
That figure gets quoted as evidence that AI is overhyped. We read it differently. Every failure pattern in that research is an organizational decision, and most of those decisions were made before anyone touched a tool.
Buying access instead of building judgment
This is the pattern we encounter most often, and it is almost invariant. We recommend that an organization invest first in upskilling its people — usually something modest and self-directed, a few hours per person. Almost nobody accepts. The response is that they would rather start using the tools now and learn as they go.
What follows is consistent enough to predict. They buy subscriptions across several platforms, frequently thousands of dollars a year, often with overlapping capabilities purchased by different departments who did not consult each other. Then four things happen.
- 01Proprietary material goes into models that may train on it. Teams upload intellectual property and internal documents without checking whether the vendor uses inputs for training. The setting usually exists. It is almost never examined.
- 02The subscription tier is chosen wrong. Either an expensive plan where consumption runs unmonitored and the bill grows without anyone owning it, or an inadequate plan producing weak output that gets blamed on AI rather than on the plan.
- 03Nobody knows the models fabricate things. Hallucination is not understood as a normal failure mode of the technology, so it is not anticipated, and confident wrong answers are taken at face value.
- 04Output is not verified. Staff do not review the content or check what came back before using it. This is the failure that compounds most quietly, because the errors surface weeks later attached to something that has already gone out.
From our engagements
Every one of those four is a training and governance failure wearing a technology costume. The organization concludes that AI did not work for them. What actually happened is that they spent the training budget on licenses and then operated the licenses without the training.
Access is cheap and immediate. Judgment is cheap and slow. Most organizations buy the first and skip the second, then wonder why the results are poor.
Choosing the visible use case over the valuable one
MIT's data shows roughly 50% to 70% of AI budgets in their executive sample flowing to sales and marketing initiatives, while the clearer cost savings were emerging in back-office functions — procurement, finance, operations.
The reason is not stupidity, it is legibility. A tool that drafts outreach or answers customer questions is easy to imagine and easy to demonstrate to a board. A tool that reduces the time to reconcile invoices requires understanding the invoicing process, which is harder to explain and much harder to make exciting. So the pitchable project wins the budget, and the project with a defensible return waits.
There is a second problem specific to customer-facing pilots: they fail publicly. A chatbot that frustrates people, copy that reads nothing like your brand, outreach volume that irritates prospects. Back-office failures are contained. Front-office failures reach your customers, which means the most visible category of pilot is also the one where a mistake costs the most.
Automating a broken process
Technology does not fix misalignment; it amplifies it. Automating a flawed process means performing the wrong work faster and more consistently than before.
This surfaces in a specific way. An organization asks for AI to accelerate a particular process. Mapping the process reveals that the actual constraint is elsewhere — the data going in is inconsistent, or three teams follow different versions of the procedure, or a step exists solely because it always has. Automating on top of that produces a faster version of a problem, and because the output now arrives with machine authority, the underlying inconsistency becomes harder to notice rather than easier.
Sometimes the honest finding is that the process needs fixing and AI is not the relevant tool. That conclusion tends to be unwelcome, and it is frequently the most valuable thing an assessment produces.
Never crossing the integration threshold
The sharpest structural finding in MIT's research is the gap between exploration and production. Over 80% of organizations had explored or piloted general-purpose assistants and roughly 40% reported deploying them. But embedded, workflow-specific tools — the kind that live inside the systems where work actually happens — reached production about 5% of the time.
That gap is the whole story. AI sitting beside your business, in a separate tab, requiring someone to remember to use it and to copy results back into a real system, is a demonstration. It generates activity that looks like progress and produces no durable change, because the moment attention moves elsewhere, usage decays.
The pilots that survive are integrated into the systems people already work in, which is unglamorous, involves permissions and data plumbing, and is the difference between a pilot and a capability.
Running a pilot that could actually succeed
Most failed pilots were not run badly. They were defined so loosely that no result could have counted as success. Four requirements fix most of that, and they are all cheap.
| Requirement | What it looks like | What it replaces |
|---|---|---|
| One named business outcome | Cut the time to produce the monthly client report from six hours to two. | "Explore how AI could help the reporting team." |
| A baseline measured first | Time the current process for three cycles before you change anything. | Estimating the before-state afterwards, which always flatters the result. |
| An owner who does the work | The person whose week improves runs it and reports on it. | A sponsor two levels removed who sees a summary at the end. |
| A defined end date and decision | Six weeks, then a documented choice: expand, adjust, or stop. | An indefinite pilot that quietly becomes permanent without ever being evaluated. |
The fourth is where discipline usually fails. Deciding to stop is treated as an admission of error, so pilots are neither expanded nor killed — they persist in a state where they consume licenses and attention while proving nothing. A pilot that ends in a clear decision to stop has done its job.
On outside help, stated plainly
The same MIT research found that externally partnered deployments reached production about twice as often as internal builds — roughly 67% against 33%. We are a consultancy, so we are the last people whose interpretation of that statistic you should accept uncritically. Take the number, note the source, and weigh it yourself.
What we would say is that the gap is unlikely to be about intelligence. Internal teams understand the business far better than any outside party will. What they usually lack is repetition — having watched the same integration fail in four different organizations and knowing which question to ask in week one. The effective arrangement pairs internal knowledge with external mileage, and where the mileage comes from matters less than that someone in the room has it.
What the 95% figure actually tells you
It is not evidence that AI does not work. It is evidence that the sequence most organizations follow does not work: buy first, learn later, choose the demonstrable use case, skip the integration, and never define what success would have looked like.
Reversed, the same list is a workable method. Establish literacy before you expand tooling. Look for value where the work is tedious rather than where the demo is impressive. Fix the process before automating it. Integrate into real systems. Define one outcome you can measure, and be willing to stop.
Stay in the loop
New writing, when there is something worth sending
Occasional notes on AI governance, adoption, and what we are seeing in client work. No newsletter cadence, no sequence — we write when we have something useful.
Next step
If a pilot already went nowhere, that is useful information.
A failed pilot usually tells you something specific about the process, the data, or the sequence. We are happy to look at what happened and tell you whether it is worth another attempt.
Start an AI Conversation