Insights

Strategy

What AI cannot do for your business yet

Knowing where these tools fall short is more useful than another list of possibilities. Three categories of work reliably disappoint, and one of them disappoints for a reason that has nothing to do with capability.

The VIP IT AI Practice8 min read

The short answer

AI underdelivers most reliably in three places: work requiring accountability nobody will delegate, tasks where the cost of a single undetected error exceeds the cumulative time saved, and processes too irregular to justify the cost of automating.

The accountability limit is not a capability gap and will not be solved by a better model. Someone has to be answerable for the outcome, and that person will keep doing the work.

Verification cost is the most commonly ignored factor. If checking the output takes as long as producing it, the tool has moved the work rather than reduced it.

These boundaries move, but not uniformly. Judge a use case on the error cost and the verification burden, not on whether a demonstration looked impressive.

Most writing about AI in business is about what is now possible. That has value, but it is not what a leader deciding where to spend the next quarter actually needs. The more useful question is narrower: where do these tools reliably disappoint, so you can avoid discovering it yourself at your own expense.

We are an AI practice, so it is fair to note the incentive here runs the other way. We are writing this because the fastest way to lose a client's confidence is to let them spend six weeks on something we could have told them would not work.

Work nobody will take accountability for

This is the most misread limitation, because it is not a capability limitation at all. There is a category of work where the deliverable is not really the artifact — it is someone's willingness to stand behind it.

A partner signing an opinion. A clinician making a call. A controller certifying figures. A principal committing to an estimate a client will rely on. In each case the output could be drafted competently by a machine, and in each case the value being purchased is that a named human is answerable if it is wrong. Removing the human does not make the process more efficient; it removes the thing being bought.

When accountability is the deliverable, a better model does not help. The person who is answerable will read it, check it, and own it — and that is where the hours go.

The practical consequence is that AI's contribution in this category is bounded by the review, not the drafting. If the reviewer must read the material closely enough to stake their name on it, a faster draft yields far less than the demonstration suggested. This is why AI-assisted work in professional services often produces smaller gains than expected — the drafting was never the bottleneck.

Where it does help is at the edges of that work: assembling the material the reviewer needs, surfacing the clause that warrants attention, drafting the routine sections so judgment can be spent on the parts that require it.

Work where one undetected error costs more than the time saved

This is arithmetic that almost nobody performs before adopting a tool. Estimate the time saved per use. Estimate the probability of an error that gets through unnoticed. Estimate what that error costs when it surfaces — including the cost of finding it, correcting downstream artifacts, and the relationship damage if a client found it first. Then compare.

For low-stakes internal work, the arithmetic is overwhelmingly favorable and you should not overthink it. For anything a client relies on, anything with a regulatory dimension, and anything feeding a decision that is hard to reverse, it frequently is not — particularly because these systems fail in a specific and awkward way. They do not signal uncertainty. A fabricated figure arrives with exactly the same fluency as a correct one, which means errors are not merely possible; they are hard to spot.

The related trap is verification cost. If output must be checked line by line to be trusted, the work has been moved rather than removed — and moved to a less pleasant form, because reviewing someone else's plausible-looking draft is slower and more error-prone than writing it yourself. A useful test before adopting anything: how long does it take to be confident this output is right? If the answer approaches the time to produce it unaided, the tool is not the answer.

TaskTime savedCost of a missed errorVerdict
Summarizing an internal meeting20–30 minutesSomeone re-reads the threadClearly worth it
Drafting a first-pass job description30–45 minutesAn awkward edit before postingClearly worth it
Producing figures for a client deliverable1–2 hoursCredibility, possibly the relationshipOnly with full verification, which usually erases the saving
Interpreting a regulatory obligationSeveral hoursPenalties and remediationResearch assistance only, never the conclusion
Extracting terms from contracts at volumeSubstantialDepends entirely on what the term governsViable with sampling and human review of exceptions

Work too irregular to be worth automating

Automation economics depend on repetition and stability. A process that runs hundreds of times a month in a consistent shape is a good candidate. A process that runs eleven times a year and looks slightly different each time is usually not, regardless of how tedious it is when it happens.

The reason is maintenance rather than build cost. Automated workflows require upkeep: systems change, a vendor alters a format, an edge case appears that the original design never anticipated. For a high-frequency process that upkeep is trivially justified. For a low-frequency, high-variance one, the maintenance can exceed the manual effort within a year, and it arrives as a series of small interruptions rather than one visible cost, which makes it easy to underestimate.

This category is where enthusiasm most often outruns judgment, because the tedious annual process is memorably painful and therefore feels like an obvious target. Frequency, not annoyance, is the criterion.

Two things it is genuinely good at, for contrast

It would be dishonest to present only the limits. Two categories consistently deliver more than people expect, and both are unglamorous.

The first is reducing the cost of starting. A blank page is expensive in a way that is rarely accounted for. A mediocre first draft that gets rewritten is often worth more than a good draft that arrives late, because it converts a task people avoid into one they can begin.

The second is retrieval across scattered information — finding the relevant thing among documents, threads, and systems where nobody remembers what exists. This is where we see the most reliable value in mid-sized organizations, and it is almost never what people ask for first.

These boundaries move, but not evenly

Everything above describes the current state, and parts of it will age. Capability improves quickly, so specific tasks will cross from the second category into the viable one, and the honest position is that we do not know how fast.

The accountability limit is different in kind. It is not waiting on a technical advance, because it is not a technical constraint — it is a question of who is answerable, and that is decided by professional obligation, liability, and client expectation rather than by capability. When those change, they will change slowly.

The practical guidance is to judge each use case on two questions rather than on whether a demonstration looked impressive: what does a single undetected error cost, and how long does it take to be confident the output is correct. Those two answers will tell you more than any feature list.

Stay in the loop

New writing, when there is something worth sending

Occasional notes on AI governance, adoption, and what we are seeing in client work. No newsletter cadence, no sequence — we write when we have something useful.

Next step

Not sure which category your idea falls into?

Describe what you are considering and we will tell you honestly whether it is a good candidate. Sometimes the answer is that it is not, or not yet.

Start an AI Conversation