Skip to content
Vlad KozakAug 5, 2026, 10:01:58 AM9 min read

Why AI pilots fail when they meet real data

 

You have probably seen the number by now. MIT's NANDA initiative reported that ninety-five per cent of enterprise generative AI pilots delivered no measurable return, a finding that travelled further and faster than almost anything else written about AI in business.

Let me be straight about that statistic before I use it, because most people quoting it are not.

It is a narrower finding than the headline suggests. The study defined success as deployment beyond pilot stage with measurable KPIs and a profit and loss impact evident within six months. That is a demanding bar, and a thoughtful critique from the Marketing AI Institute points out it excludes efficiency gains, cost avoidance, churn reduction and pipeline effects, none of which show up cleanly in a P&L inside two quarters. It also rests on a modest interview base that the authors themselves describe as directionally accurate. Anyone waving it around as proof that AI does not work is overreaching.

So I do not think ninety-five per cent of AI projects are worthless. I do think the direction of travel is real, and it matches what I see. Gartner reached a similar conclusion from a different angle, predicting that more than forty per cent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Their analyst described most current projects as <cite index="22-1">"mostly driven by hype and are often misapplied"</cite>. Not incapable. Misapplied.

Here is what I want to add to that, from watching it happen up close.

Pilots almost never fail at the model. They fail at the seam between the demo and the real thing.

The moment it goes wrong

The pattern is consistent enough that I can now predict it.

Someone builds a prototype. It works. It works genuinely well, in the room, in front of the people who commissioned it. Everybody is pleased. A production budget gets discussed with real enthusiasm.

Then the thing meets the business, and within a fortnight the same three sentences get said in some order. That is not how we actually do it. It has not accounted for X. That number looks wrong.

Nothing has broken. The model is the same model. What has changed is that the input is no longer the version prepared for the demonstration. It is the real version, with the merged supplier records and the branch that does things differently and the rule nobody wrote down because everybody in the room already knew it.

This is not a technology failure and it is not incompetence. It is a scoping failure, and it is almost always avoidable.

Four ways real data kills a promising pilot

The data was cleaned for the demo, and nobody said so out loud.

Rarely deliberate. Somebody pulls an extract for the prototype and, because they are conscientious, they tidy it. Duplicates removed, obvious errors corrected, the odd date fixed. That extract is now a description of a business that does not exist. Every judgement made about the prototype's accuracy is a judgement about an artificial dataset.

The fix is uncomfortable and simple: build on the unmodified production data from the first day. It will look worse. It will be true.

Nobody had written down how the decision is actually made.

This is the big one and it is the most consistently underestimated.

Ask an experienced buyer how they decide reorder quantities and you will get four or five factors. Watch them do it for three weeks and there will be fifteen, several of which they have never articulated, at least two of which contradict each other depending on the supplier, and one of which is that they do not trust a particular lead time and quietly pad it.

None of that is in a specification. None of it is in the system. It is in a person, and it is real business logic that has been keeping your margins where they are. A pilot built from the four factors they can articulate will produce recommendations that are technically defensible and obviously wrong to the person who knows.

The edge cases turn out to be the job.

In a demo, edge cases are a rounding error. In an operation, they are frequently where the money is. The unusual supplier arrangement, the customer with bespoke pricing, the stock item that behaves differently because of how it is packed. A pilot that handles the ordinary ninety per cent and falls over on the exceptions has automated the easy part and left the expensive part untouched.

The system is right, and nobody believes it.

This one is genuinely underrated. You can build something accurate that fails anyway, because the people who would have to act on it cannot see why it is saying what it is saying.

If a recommendation arrives with no reasoning and no confidence signal, an experienced person has two options: accept it on faith or ignore it. Most sensibly ignore it. Trust is not a soft factor here, it is a delivery requirement, and it is earned by showing the working.

The test I would insist on

There is a way to find all four of these problems for very little money, and I would now refuse to run a pilot without it.

Run it in parallel against decisions your people were going to make anyway.

Not a test dataset. Not a retrospective backtest. Live decisions, in the normal weekly rhythm, with your team using the prototype alongside their existing process and nobody obliged to follow it.

Then look only at the disagreements.

Every disagreement between the system and the experienced human is information, and it is exactly one of three things:

  • A gap in the data. The system did not know something it needed to know. Now you know what to connect next.
  • A rule you have not captured. The human is applying judgement that has never been written down. Now you can write it down, and that document is valuable whether or not the AI project proceeds.
  • A better answer. Sometimes the system is right and the habit is wrong. These are the ones that pay for the work, and you will not find them any other way.

A few weeks of that is worth more than any amount of specification, because it produces evidence rather than assumptions. It also has an effect nobody plans for: the people who were sceptical become the people who improve it, because they are the ones finding the flaws and being listened to.

When we recommended against going to production

We were working with a multi-store specialty retailer. Inventory is the largest asset on their balance sheet and it turns roughly once a year, so the money at stake is not marginal. After a Discovery Workshop and a series of working sessions, inventory intelligence came out as the highest-value use case, comfortably.

The plan was to move from workshop straight to a production build. That is what the client wanted and it is what we had scoped.

Partway through, we recommended against it.

Not because the approach was wrong. Because real questions were still open about lead times, about outstanding orders, and about how much to trust one upstream system. We could have built to specification, hit every stated requirement, and delivered something that produced recommendations the buying team would have found reasons to distrust within a month.

So instead we deployed the working prototype online for the team to use on real weekly reorder and redistribution decisions, comparing its recommendations against what the buyers would have done anyway. The buying rules that had lived only in one person's head got written down, by her, in plain language, and encoded.

That is a smaller invoice than the one we had scoped. It is also the reason the production build will be scoped from evidence rather than assumption, which is worth considerably more to them than a few weeks of saved time.

I mention it because "we recommended against the bigger piece of work" is the sort of thing consultancies say and rarely do. The test of whether a firm means it is whether they have ever sent the smaller invoice.

What a pilot should have before it starts

If you are about to commission one, I would want all six of these agreed in writing beforehand.

A named decision it is meant to improve. Not a capability. A decision somebody makes, on a known cadence, that is currently slow or wrong.

Success criteria you could argue about. Defined before anyone builds, specific enough that reasonable people could disagree about whether they were met. "Improved visibility" is not a success criterion.

Production data, unmodified. Written into the scope, so nobody helpfully tidies it.

A real evaluation period with real users. Weeks, not a demonstration. Running against live decisions.

A fixed fee. A pilot on time and materials has no natural end, which is how you get the pilot that quietly never finishes. Gartner's cancellation reasons put escalating cost first for good reason.

An agreed answer to "what if it does not work". Ask any consultant this before you engage them and listen carefully to the reply. If there is no cheap way to find out, and no version of the engagement where they tell you to stop, you are not buying a pilot. You are buying a build with an optimistic name.

Not a model problem

The New Zealand picture supports this reading. A Publicis Sapient survey of AI decision-makers found forty-two per cent saying their organisation is not set up to capture the value even while most reported using AI regularly, with operating structure rather than technology identified as the obstacle. Xero's New Zealand research found strong SME enthusiasm alongside real uncertainty about how to bring the technology into day to day operations. Stats NZ and MBIE are currently running a national survey of around 20,000 businesses with findings due later this year, which should give us the first properly representative local picture.

Meanwhile the models keep getting better, which is a good thing and also a distraction. Anthropic's engineering team has written the clearest thing I have read on this, framing context as a finite resource to be curated rather than a window to be filled. Their observation that even a strong model cannot compensate for badly assembled context matches everything we see in production.

The best line I have heard on it came from a conference session: better models do not fix fractured context. If your pilot failed, it is worth asking what the model was actually given, and whether anybody had written down how your business really makes the decision you asked it to make.

Usually nobody had. That is fixable, and it is cheaper to find out in six weeks than in eighteen months.


Vlad Kozak leads Ideation Partners, a data and AI consultancy in Auckland. The practice is backed by Verde Group, whose teams have implemented and supported the systems New Zealand businesses run on for more than twenty-five years.

Related: Getting usable data out of the systems you already run

If you are deciding whether to run a pilot, or working out why the last one stalled, take a free hour. We will tell you honestly whether there is anything worth doing, and say so if there is not.

avatar
Vlad Kozak
Our illustrious leader embodies the optimisers’ culture: part programmer, part solution designer, part sales evangelist, 100% solution focused. Vlad started with CBA straight from University, did stints with Oracle, Epicor and another ERP reseller, before becoming a founding shareholder and executive at Greentree. There at the start, Vlad is responsible for many of the key features that today differentiate Verde software from its competitors.

RELATED ARTICLES