AI ROI measurement

AI ROI Measurement: Metrics That Survive Scrutiny

September 8, 2026 · 7 min read · Autana Solutions, Vancouver
AI ROI Measurement: Metrics That Survive Scrutiny — Autana Solutions

Most AI projects end with a meeting where somebody says it's working. The hard part is proving it.

That gap matters more than people admit. AI ROI measurement is the difference between a tool you keep and a subscription nobody remembers approving. The research on this is unusually clear, and most of it cuts against optimism. If you measure honestly, you will sometimes learn the automation didn't help. That's the point. Finding out in month two is cheap. Finding out in year two is not.

Your gut is a terrible instrument

Start with the most uncomfortable finding in the field. A randomized controlled trial by METR researchers had 16 experienced open source developers complete 246 real tasks on codebases where they averaged five years of prior experience. Before the study, the developers forecast that AI tools would cut their completion time by 24%. Afterwards, they estimated they had been sped up by about 20%. The measured result was the opposite: allowing AI tools increased completion time by 19%.

Sit with that for a second. These are skilled people, working on code they know well, and their felt sense of speed was off by roughly 39 percentage points in the wrong direction. Experts asked to forecast the outcome were wrong too.

A second study points at why. Researchers surveyed 415 software practitioners using the SPACE framework and found that reported speed gains often hid a shuffle rather than a saving: heavier code review, ongoing cognitive load from verifying output, and collaboration patterns that didn't change. Their conclusion is that perceived productivity gains may be spurious, surface level acceleration accompanied by redistributed effort and hidden costs.

Both studies look at software developers, which is a narrower setting than a plumbing dispatcher or a dental clinic front desk. But the mechanism travels. Work that feels faster and work that is faster are separate things, and only one of them shows up on a bank statement.

The metrics that hold up

A metric survives scrutiny when it is countable without anybody's opinion, tied to a unit of work you already care about, and measured the same way before and after. That's it. Six that qualify:

  • Cycle time on one named unit of work. Minutes from missed call to callback. Hours from quote request to quote sent. Days from job complete to invoice paid. Pick the unit before you deploy anything.
  • Autonomous completion rate. The share of interactions the system finished with nobody stepping in. Track its twin, the escalation rate, in the same report.
  • Rework rate. Corrections, callbacks, refunds, re-dos, and apology emails. This is where the hidden cost from the SPACE study shows up if it exists.
  • Fully loaded cost per completed unit. Software plus usage fees plus the staff time spent supervising, correcting, and prompting.
  • A holdout. Leave one location, one shift, or one lead source unautomated for eight weeks. Comparing to your own untouched baseline beats comparing to a memory.
  • One downstream business number. Appointments actually booked and attended. Invoices paid within 30 days. Quotes accepted. If none of these move, the operational wins may not be reaching the ledger.

The vanity list is shorter and easier to spot. Messages handled, seats licensed, hours saved as reported by the people using the tool, demo quality, and public benchmark scores for whichever model sits underneath. Worst of the bunch is the estimated hours saved multiplied by an hourly rate. That calculation is built from exactly the self report the METR trial showed to be unreliable, and it always produces a flattering number.

Write the baseline down before you switch anything on

The most valuable measurement work happens before deployment, because after go live your baseline is gone. You cannot reconstruct last quarter's average response time from feelings.

The NIST AI Risk Management Framework 1.0, published January 26, 2023, organizes this into four functions: Govern, Map, Measure and Manage. The Measure function calls for quantitative, qualitative or mixed methods to analyze, assess, benchmark and monitor AI risk. Two subcategories are worth stealing outright even if you never adopt the rest. Measure 2.1 asks that test sets, metrics, and the tools used be documented. Measure 2.4 asks that the system's functionality and behaviour be monitored once it's in production, not just at launch. The framework is free, vendor neutral, and a better checklist than anything a software salesperson will hand you.

Put a real number on the cost side

ROI has a denominator, and vendors are vague about it. Per token pricing is public, so use it.

Checked against Anthropic's official pricing documentation in September 2026, Claude Haiku 4.5 is published at $1 per million input tokens and $5 per million output tokens. The same page works a support example: at roughly 3,700 tokens per conversation, 10,000 support tickets come to about $37 in model cost. The docs also list a 50% discount for batch processing and cache reads at 0.1x the base input price for that model tier.

Thirty seven dollars per ten thousand tickets is not what makes or breaks a deployment. That's the useful lesson. Model usage is usually the smallest line in the budget. Integration work, supervision time, and rework are the big ones, and they're the ones nobody puts in the spreadsheet.

What the Canadian numbers actually say

Statistics Canada reported that in the second quarter of 2026, 19.2% of Canadian businesses used AI to produce goods or deliver services over the preceding 12 months, up from 6.1% two years earlier. Virtual agents and chatbots were in use at 28.2% of AI using businesses. Notably, 44.4% reported changing their training or staffing practices because of AI, and about 40% said AI use was simply irrelevant to their operations.

The harder finding comes from a Statistics Canada study by Jiang Li and Huju Liu, released April 22, 2026. AI adopting firms initially showed a 16.8% labour productivity premium. Adjusting for pre-existing differences dropped it to 10.2%. Once complementary investments were accounted for, it fell to 5.1% and became statistically insignificant. The authors conclude there's no statistically significant direct association between AI adoption and productivity once proper controls are applied.

Read that as encouragement to do the surrounding work rather than as a reason to skip AI. Adoption alone isn't the win. The data plumbing, the process redesign and the training are where the measurable gains seem to live. A Burnaby shop that bolts a chatbot onto an unchanged intake process is buying the adoption without the complements.

Where this doesn't apply

Some honest limits.

If your volume is low, skip the ROI math. Under a few hundred repetitions a month, normal variation swamps any effect you're trying to detect, and an eight week comparison will tell you nothing reliable. Judge quality and error rate instead, and revisit the economics later.

The developer studies are about experienced engineers on mature codebases. They're a warning about self reported gains, not a prediction about your after hours phone agent. Don't quote the 19% figure as if it applies to receptionist work.

Some wins resist pricing. Never missing a 9pm call has real value that shows up months later, scattered across jobs you'd never have known you lost. Say plainly that you're accepting it on judgment rather than dressing it up in a fake number.

And if you have no baseline at all, the first month of the project should be spent measuring the manual process. That feels like a delay. It's the only thing that makes month six interpretable.

Where error costs are high, hiring, credit, health, safety, the calculation changes shape entirely. There the question isn't how much time you saved. It's how often the system is wrong and what happens when it is.

Sources

  • Becker, J., Rush, N., Barnes, E., and Rein, D. (2025). *Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity*. arXiv:2507.09089. https://arxiv.org/abs/2507.09089
  • Afroz, S., Feng, Z., Menezes, T., Kimura, K., Trinkenreich, B., Steinmacher, I., and Sarma, A. (2025). *The Fast and Spurious: Developer Productivity with GenAI*. arXiv:2510.24265. https://arxiv.org/abs/2510.24265
  • National Institute of Standards and Technology (2023). *AI Risk Management Framework 1.0*. https://www.nist.gov/itl/ai-risk-management-framework
  • Statistics Canada (2026). *Analysis on artificial intelligence use by businesses in Canada, second quarter of 2026*, released June 11, 2026. https://www150.statcan.gc.ca/n1/pub/11-621-m/11-621-m2026010-eng.htm
  • Li, J. and Liu, H., Statistics Canada (2026). *Artificial intelligence adoption and productivity in Canadian firms*, April 22, 2026. https://www150.statcan.gc.ca/n1/pub/36-28-0001/2026004/article/00002-eng.htm
  • Anthropic. *Pricing*, Claude platform documentation, accessed September 2026. https://platform.claude.com/docs/en/about-claude/pricing

If you're weighing an AI deployment across Metro Vancouver, or you already have one running and can't tell whether it's earning its keep, we're happy to walk through the baseline with you. Autana Solutions is based in Burnaby and New Westminster, and the first call is free. We'd rather help you pick a metric you can defend than sell you something you can't measure.

AI ROImeasurementautomationsmall businessVancouver

Want an AI employee for your business?

We install a 24/7 AI worker for businesses in Vancouver, Burnaby, and beyond. Book a free Discovery Call.

Book a call

Keep reading