Measuring AI Projects: Don’t Fall for the Vanity Metric Trap 

A software project can meet its stated objective and still fail to deliver meaningful business value. 

The system works. The automation rate is achieved. Users adopt it. On paper, the project is a success. 

Yet when the business looks for the promised outcome – time saved, cost reduced, capacity released – the impact is often far smaller than expected. 

This is the vanity metric trap: measuring what the system does, not what the business actually needs to change. 

It is becoming more common as AI expands what we can automate. We can now process language, interpret variation, and handle ambiguity that previously required human judgement. But capability does not guarantee value. 

A project can hit its target and still miss the point entirely. 

Automating 70% of Orders Did Not Save 70% of the Work 

We worked with a manufacturer processing more than 50 customer orders per day via email. The workload varied significantly – from simple one-line requests to complex orders containing 30 or more line items. 

Three full-time employees managed the process. The goal was to free up enough capacity for one person to move into another area of the business. 

The solution used AI to extract order details from emails, match products and customers against the ERP system, and generate sales orders automatically. This was necessary because customers rarely used consistent product codes or terminology, often relying on shorthand or partial descriptions. 

The target was simple: automate 70% of inbound orders. 

The project achieved it. 

But the expected time saving never materialised. 

The reason became clear only when we looked beyond the headline metric. The 70% that was automated consisted mostly of the simplest orders – the ones that already took one to five minutes to process. The complex, time-consuming orders, which drove most of the effort, remained largely untouched. 

The metric was technically correct, but structurally misleading. It treated all orders as equal when they were not. 

A more meaningful measure would have been the percentage of line items automated, total staff hours removed from the process, or reduction in average handling time. Any of these would have shifted focus toward where the real cost sat. 

This is a common pattern we see in early-stage automation programmes. Once you step back from the stated target and examine how work is actually distributed, the real opportunity becomes visible. In this case, the organisation wasn’t failing to automate enough – it was measuring the wrong thing entirely. 

The system worked. The business outcome did not. 

A Process Can Be Automated and Still Become Irrelevant 

In another case, a legal practice had a structured onboarding process for new clients and matters. It was predictable, but varied enough in language and intent that AI was introduced to interpret requests and route them correctly. 

The system performed well and initially reduced administrative effort. 

However, the business itself was evolving faster than the process it had automated. 

As the firm grew, services expanded and internal ways of working shifted. The onboarding workflow that had been optimised became less central to how the organisation actually operated. 

The project had successfully automated a process the business was gradually moving away from. 

This is a subtle but important failure mode: optimising efficiency in a static view of a dynamic organisation. 

A more valuable objective would have focused on adaptability rather than throughput – how quickly new workflows could be introduced, how easily processes could be adjusted, and whether the system could evolve alongside changing service models. 

In practice, this required challenging an early assumption: that the onboarding process itself was a stable, long-term anchor. That conversation is often uncomfortable, but it is also where the most meaningful value is uncovered. 

Efficiency only matters while the process remains relevant. 

The Business Was Not as Standardised as It Believed 

A civil services organisation approached a similar challenge with a clear target: automate 80% of incoming service requests. 

On the surface, the process appeared standardised and well understood. 

In reality, only around 30% of requests followed the documented “standard” process. 

The remainder were shaped by a mix of historical exceptions, customer-specific agreements, and informal practices held in staff memory rather than in systems or documentation. 

Employees knew these variations instinctively. They understood which customers required different handling, which phrases implied specific conditions, and which exceptions had effectively become the norm. But none of this knowledge was formally captured. 

The automation could handle the 30% cleanly. The remaining 50%+ required the system to replicate years of tacit knowledge the organisation had never explicitly defined. 

At first glance, this looked like a failure to meet the target. 

In reality, it revealed something far more important: the organisation did not have a single standard process – it had a collection of negotiated ones. 

This is where deeper analysis becomes critical. The real opportunity was not simply automation, but surfacing and structuring hidden operational knowledge: identifying recurring customer patterns, exposing undocumented service rules, and distinguishing true exceptions from assumed ones. 

Only once that became visible could the business make informed decisions about what should be standardised, what should remain flexible, and what could be safely automated. 

A rigid 80% target would have completely missed this insight. 

The Trap: When the Metric Looks Right but the Outcome Is Wrong 

Vanity metrics are not usually wrong. They are simply incomplete. 

They are easy to measure, easy to report, and easy to optimise: 

  • percentage of transactions automated 
  • number of AI outputs generated 
  • system usage or login counts 
  • workflow completion volumes 

But none of these answer the real question: 

Did the business actually improve? 

A high automation rate does not guarantee reduced effort. 
A high output volume does not guarantee quality. 
A high usage rate does not guarantee value. 

The risk is not deception – it is misplaced confidence in the wrong signal. 

Measure the Outcome, Not the Activity 

Effective measurement starts with the business outcome, not system behaviour. 

If the goal is capacity, measure total time saved across the full process. If the goal is speed, measure end-to-end turnaround time. If the goal is quality, measure errors, rework, and overrides. 

And critically, measure what happens after automation is introduced. 

Work rarely disappears – it moves. Sometimes into review queues, exception handling, or downstream correction effort. A system that “automates” a task but shifts effort elsewhere is not reducing work; it is redistributing it. 

Useful measurement typically spans: 

  • Efficiency: time, cost, effort 
  • Quality: accuracy, rework, exceptions 
  • Outcome: turnaround, capacity, service levels 
  • Adoption: sustained real-world usage 
  • Flexibility: ability to evolve with the business 

Not every project needs complexity, but every project needs honesty about what is actually improving. 

Baselines Reveal What Targets Hide 

A target without a baseline is just an assumption. 

Before automation begins, it is essential to understand where time is actually spent, which cases drive disproportionate effort, how often exceptions occur, and what variation exists in real-world work. 

This does not require heavy analysis. In most cases, a small sample of real transactions reveals more than months of assumptions. 

The key is avoiding averages. Averages smooth out the very variation that creates cost, delay, and complexity. 

Across the examples in this article, averages consistently hid the truth: 

  • order volume hid effort distribution 
  • process documentation hid real-world variation 
  • workflow design hid organisational change 

Without a baseline grounded in reality, optimisation simply reinforces existing assumptions. 

The Hard Part: Challenging What the Business Thinks It Knows 

The most valuable work in any automation or AI initiative often happens before any system is built. 

It is not technical – it is interpretive. It is about understanding what is really happening inside the organisation and whether the stated problem is actually the right one to solve. 

That can surface uncomfortable truths: 

  • the “standard process” may not be standard 
  • the highest-volume activity may not be the highest-cost 
  • the most visible work may not be the most valuable 

This is where experienced partners add real value. At CROFTI, much of our work sits upstream of delivery – helping organisations uncover how work actually flows, where effort is really spent, and which assumptions are quietly shaping project success criteria. 

Because if you optimise the wrong thing, you simply do it faster. 

Successful Delivery Is Not the Final Measure 

Software still needs to work. It must be reliable, secure, and usable. 

But that is only delivery. 

Value is something else entirely. 

The real question is whether the organisation is better off than before: less effort to achieve the same outcome, faster delivery of meaningful work, improved quality and consistency, and greater capacity to adapt and grow. 

And just as importantly, whether the organisation now understands itself better than it did before the project began. 

Related reading: the same pattern shows up in delivery. AI reduces the effort of writing code, but increases the work of defining behaviour, validating outputs and earning user trust. We unpack that shift in AI Has Changed Software Development, But Not in the Way Most Businesses Expect. 

Where to Start 

Most organisations do not struggle with building software. They struggle with defining what “better” actually means. 

If you are exploring automation or AI, the most important first step is not selecting a tool or defining a feature set – it is understanding the work itself. Where effort is really going, where variation exists, and which assumptions are quietly embedded in current measures of success. 

Start by questioning the metrics you are planning to use. Ask whether they reflect activity or outcome. Then test them against real examples of work, not averages or summaries. 

This is often where the biggest shift happens – before any technology is introduced. 

At CROFTI, this is the space we work in. We help organisations move from assumed processes to observed reality, from output metrics to outcome measures, and from “what we think is happening” to what is actually driving effort and value. 

If you are at the point where a project looks successful on paper but uncertain in impact, that is usually the right moment to step back and reframe the problem before moving forward. 

Because the most important decisions are rarely about what to build – they are about what to measure in the first place. 

This thinking sits at the front row of how to run AI-First Development: agreeing what should actually improve, and how it will be measured, before anything is built. If you are not yet at that point, Discovery & Design is the step that gets you there.