AI pilots fail because they test a tool instead of changing a workflow. Nobody owns the result, nobody measured the work before, and the pilot was never designed to go live. To get one into production, turn it around: pick one workflow with an owner, measure it before you touch it, redesign the steps, keep a person checking the AI, and set a go-live date before you start.
How often do AI pilots reach daily use?
Not often. A preliminary 2025 report from MIT's Project NANDA looked at more than 300 public AI initiatives and spoke with leaders at dozens of organisations between January and June 2025. In its research, about 60% of organisations had evaluated task-specific AI tools, about 20% reached a pilot, and about 5% had one in production with a measurable effect (DX summary, Virtualization Review).
Treat those figures with care. The report is preliminary and its sample is its own, so the exact percentages will not describe your company. But the direction matches other research. McKinsey's 2025 global survey found that 88% of organisations use AI, nearly two-thirds have not started to scale it, and only 39% report any effect on EBIT (McKinsey).
The MIT authors put the cause less on the models and more on implementation: tools that do not learn from feedback, and tools that do not fit the way people already work. That matches what we hear from CEOs and owners every week.
What a stalled pilot looks like
Most stalled pilots look alike. A company gives twenty people access to an AI tool for three months. There is a kick-off, a training session and a Teams channel for tips. Usage is high in the first weeks. At the end, a survey says most people like the tool and feel more productive.
Then someone asks what changed. Which work takes less time now? Which client gets an answer faster? Nobody can say, because nobody wrote down how the work ran before, and no workflow was changed on purpose. The licences are renewed or dropped based on a feeling, and the next pilot starts the same way.
None of this is anyone's fault. A pilot that tests a tool can only answer the question it asked: do people like the tool? It cannot answer the question the board cares about: does the company work better?
Five reasons pilots stall
1. The pilot tests a tool, not a workflow
A tool spread thinly across many tasks helps a little everywhere and changes nothing in particular. People use it to draft emails, summarise meetings and rewrite documents. Each use saves a few minutes, scattered across the week, with no single process running differently.
A workflow is different. It has a start and an end, a volume and an owner. When AI takes over the checking step in client onboarding, the whole workflow runs differently, every time. That is a change you can see, measure and keep.
2. Nobody owns the result
Pilots are often owned by IT, an innovation team or an enthusiastic individual. None of them owns the work the pilot is supposed to improve. When the pilot ends, there is nobody in the business whose job is to keep the new way running.
Every change needs a workflow owner: one person in the team that does the work, who approves the new version and is accountable for the result. Without that person, a pilot is a demonstration.
3. There is no baseline
If nobody measured how long the work took before, nobody can show what changed after. The pilot ends with opinions, and opinions lose to budget pressure. A baseline takes days, not months: the steps of the workflow, how often each happens and how many minutes it takes, checked against real cases. We explain the method in how to measure AI ROI.
4. The new way does not fit the old work
Many pilots ask people to leave their normal tools, open an AI tool, paste something in, and paste the result back. That works for a week. Then the extra steps feel like friction and people drift back to the old way.
The AI has to live where the work already happens: in the inbox, the CRM, the document system, the Teams channel. And it has to learn. When a person corrects the AI, that correction should become a rule, so the same mistake does not come back. A tool that makes the same mistake every week loses trust quickly.
5. The road to production starts too late
Most pilots run on test data with no access to real systems. Only when the pilot succeeds does the real work start: security review, access to client data, approval from IT, a conversation with the works council. By then the momentum is gone.
Design for production from the first day. Agree which data the workflow may use, who approves it, which model it runs on and how people are informed. Then the pilot and the first version of the real workflow are the same thing.
A pilot and a first workflow are not the same
| A typical pilot | A first workflow | |
|---|---|---|
| Goal | Learn whether people like a tool | Change how one piece of work runs |
| Scope | A tool for a group of people | One workflow, end to end |
| Owner | IT or an innovation team | A workflow owner in the team that does the work |
| Measure | Usage and a survey | Minutes, waiting time and rework, before and after |
| Data | Test data | Real cases, with agreed access |
| Human role | Optional | A person checks what the AI does |
| End | A report | A go-live date |
How to get one workflow into production
This is the order we use. The first workflow typically runs on live work within eight weeks, including the baseline.
- Choose one workflow with an owner and a number. Pick work that repeats many times a week and costs hours. Name the owner and the number you want to move.
- Measure it first. Write down the steps, the minutes and the volume. Check them against real cases. Date the baseline.
- Design the new version. Mark every step as automated, AI with a human check, or a tool or instruction. Keep a person on the steps that carry judgement or risk.
- Settle the rules up front. Data access, approvals, the model and where it runs, and how the team and the works council are informed.
- Build it where people work. In the inbox, the CRM, the document system. Not in a separate tool people must remember to open.
- Test on past cases. Run the new version on 20 recent cases and compare the results with what people actually did.
- Go live small, then wider. One or two people first, with a person checking every output. Then the whole team.
- Turn every mistake into a rule. When a person corrects the AI, write the correction down so it does not happen again.
- Measure again. The same way as the baseline, after a few weeks of live work. Let finance confirm the result.
The details of each step are in how we work.
What the successful minority does differently
The organisations that get results from AI are not the ones with the most tools. McKinsey describes a small group of high performers, around 6% of its sample, that attribute more than 5% of EBIT to AI. What sets them apart is not spending alone: they redesign workflows, leadership owns the change, and they keep people in the loop where it matters (McKinsey).
The MIT research says something similar from the other side: the minority that succeeds builds AI into the work, and uses tools that learn from feedback. The common thread is that both treat AI as a change in how work runs, not as software to roll out.
What to do with the pilots you already run
Most companies have a few pilots in flight. Look at each one and sort it into one of three groups.
Keep it if it already changes one workflow, has an owner in the team and has numbers before and after. Make it permanent and measure it again in three months.
Convert it if the tool is good but no workflow changed. Pick the one workflow where the tool helps most, name an owner, take a baseline and redesign that workflow around it.
Stop it if nobody owns it, nobody uses it after the first weeks and nobody can name the work it should improve. Stopping a pilot is not a failure. Keeping a pilot that nobody owns is.
Frequently asked questions
How long should an AI pilot run?
Long enough to measure live work, short enough to keep momentum. A first workflow should run on real cases within about eight weeks. If nothing is in daily use after twelve weeks, change the approach or stop.
Should we build our own AI tools or buy them?
Start with what you already have: Microsoft 365 or Google Workspace, your CRM, your accounting system. Build only what connects them for a specific workflow. The goal is a changed workflow, not a new piece of software.
Who should run the first workflow?
The workflow owner in the team that does the work, supported by one person who owns AI across the company. IT connects the systems; it should not own the result.
What if people resist the change?
Involve them from the start. The people who do the work design the new version with you, test it and decide when it is ready. Resistance usually comes from changes done to people, not with them.
Have pilots that went nowhere? Book a 30-minute call and we will look at which one to turn into a workflow. For the full approach, read where to start with AI in a mid-size company.