Organisations rarely fail at AI because the technology disappointed them. They fail because they automated the wrong process, and could not tell whether it worked.
The technology is now capable enough that selection matters far more than capability.
Start with hours, not with hype
Build a list of every recurring task and estimate the hours it consumes monthly. Not what people say the job involves — what they actually spend time on, measured over a fortnight.
The results are usually surprising. The tasks executives assume are expensive are often cheap, and something nobody mentions turns out to consume forty hours a month across three people.
Rank by hours multiplied by the cost of that hour. That list, not a vendor demo, is your automation pipeline.
The four filters
A good first AI or automation candidate passes all four:
High volume, high repetition. Something done hundreds of times beats something done well once. Automation return scales with frequency.
Clear inputs and outputs. If you cannot describe the transformation precisely, you cannot verify the result, and you will not know when it is wrong.
Tolerable error cost. Start where a mistake is inconvenient rather than dangerous. Invoice coding errors are recoverable; a clinical dosage error is not.
Measurable baseline. If you cannot state the current cost, time or error rate in a number, you cannot prove improvement — which means you cannot justify the next project.
Candidates that fail the fourth filter are the most common trap. They produce impressive demos and unprovable results.
Fix the data before the model
This is where timelines are actually decided. A forecasting project on clean, consolidated data is a few weeks. The same project on data scattered across four systems with inconsistent categories is a quarter, and most of that quarter is data work.
Be honest about this in planning. Teams that promise a model in six weeks on data that needs three months of consolidation miss both dates.
The useful reframe: the data consolidation is valuable even if the model never ships. It is rarely wasted work.
Explainability is a requirement, not a luxury
A recommendation nobody can interrogate will not be used, however accurate it is. Buyers ignore forecasts they cannot question. Managers override classifications they cannot understand. Within three months the system becomes decoration.
Insist that every automated output shows its drivers: which inputs mattered, what the confidence is, and what a human can do to challenge it.
Adoption follows trust, and trust follows transparency.
Keep a human in the loop until you have earned the right to remove them
The pattern that works:
- Observe — the model runs alongside the existing process, and you compare outputs for a period without changing anything
- Advise — the model recommends, a human approves every case, and you track override rates
- Act with review — the model acts on high-confidence cases, a human reviews exceptions
- Act with sampling — the model acts autonomously, with periodic sampling and drift monitoring
Most organisations want to start at stage four. The ones that get there sustainably pass through the first three and can show you the data from each.
Override rate at stage two is the metric that matters. If humans override 40% of recommendations, you do not have an automation, you have an extra step.
Monitor from day one
Models degrade. Data distributions shift, source systems change formats, seasonality behaves unexpectedly, and a category that was rare becomes common.
Put in place, at launch rather than later:
- accuracy tracking against actual outcomes with a defined lag
- drift detection on input distributions
- alerts when confidence drops or error rates rise past a threshold
- a scheduled retraining cycle with a named owner
- a documented rollback to the previous process
A model without monitoring is a liability that produces confident wrong answers at scale.
What to avoid as a first project
Some categories consistently disappoint as first attempts:
- Open-ended customer conversation without a defined handoff to a human. Errors are public and reputational.
- Anything safety-critical or clinically consequential. The verification burden dwarfs the benefit at small scale.
- Processes with no stable definition. If the process changes monthly, you are automating a moving target.
- Projects chosen because a competitor announced one. Their process is not your process and their baseline is not your baseline.
A realistic first-year plan
Quarter one: measure hours, build the ranked list, consolidate data for the top candidate.
Quarter two: ship the first automation in observe-then-advise mode, with measurement in place.
Quarter three: move it to act-with-review, publish the measured saving, select candidate two.
Quarter four: ship candidate two, review the first one's sustained performance, decide whether to expand.
Two proven automations with documented returns in year one beats six pilots with none. It also means the second year gets funded, which is the actual measure of success.
Applying this to your own organisation?
We would rather diagnose your situation than sell you a product. Book a call and we will tell you honestly whether this is worth doing now, later, or not at all.