Every vendor in AI has an ROI story. Most are wrong in the same way.
“Customers save 10 hours per week with AI.” “Teams see a 40% productivity boost in month one.” “Payback in under three months.” These numbers confuse activity with outcome, pilot with scaled rollout, vendor benchmark with your context, and AI-specific gains with gains that would have happened anyway.
This post is for leaders who need to know whether AI is actually working — whether the work produces real value that wouldn’t exist without it. The honest answer is usually smaller than the vendor’s number, and the measurement system that produces it is more useful than the bigger number that’s not credible.
The AI ROI claims that don’t survive contact with reality
Four patterns make most AI ROI claims dishonest:
Activity conflated with outcome. “Users submitted 50,000 prompts last quarter.” So what? Prompt count is not value. The right question is how many of those prompts changed a downstream decision, and what that decision was worth.
Pilot results projected as steady state. A pilot is six engineers with a curated workflow and a champion in the room. The “30% productivity gain” is real for those six engineers in those weeks — and not real when the same approach is spread across 600 people in 40 teams with no per-team support. Scaling AI destroys most of the gain unless you’ve designed for scale.
Vendor benchmarks applied to your work. A vendor’s benchmark on a 200-task evaluation suite is not your work. The benchmark measures general capability; your work has specific inputs, tolerance for error, and time costs. The vendor’s number is a ceiling on tasks the vendor selected. Plan for less.
Attribution without a counterfactual. “Our team is 25% faster since adopting AI.” Was the team faster because of AI, or because you also hired two new people and refactored the workflow? Most AI ROI claims cannot answer the counterfactual question. Without an answer, the number is a guess.
Hard value vs soft value: the first distinction
The cleanest framing for AI ROI measurement is two categories.
Hard value is value you can put on a financial statement. Time saved on a paid task. Error rate dropped where errors had a known cost. Revenue influenced by AI-assisted work where the influence is traceable. Capacity unlocked that converts to new business or deferred headcount. This is what the CFO will accept.
Soft value is real but doesn’t appear on a financial statement. Decisions made better. Stress reduced. Faster iteration cycles. Higher quality output that doesn’t translate to a measurable revenue line. Soft value is often larger than hard value in the first 18 months — and harder to defend.
The mistake is to claim soft value as hard value (most vendor ROI pitches do this), or to ignore soft value entirely (most CFO-driven programs do this). The discipline is to measure both, separately, with different methods, and report them as different numbers.
For the structured framework that separates hard from soft value and ties them to specific business outcomes — use case identification, ROI attribution, sensitivity analysis — AI in Business — Level 4 has the measurement module that goes deeper.
Leading indicators vs lagging indicators
The second distinction is timing.
Lagging indicators tell you if AI worked, six to twelve months after the fact. Revenue influenced, cost reduced, time saved on the bottom-line task. These are what most people mean by “ROI,” and the only ones a CFO will trust.
Leading indicators tell you if AI is likely to work, weeks or months before the lagging indicator resolves. Prompt reuse rate among trained teams. Workflow adoption depth — are people using AI for the three tasks you targeted, or for everything except those three? Output quality spot-check scores. Leading indicators are how you steer before you’ve wasted a year on the wrong approach.
The mistake is to skip leading indicators and wait for lagging ones. By the time a lagging indicator resolves, the rollout has either succeeded or failed — you’re reading the scoreboard instead of changing the game. The discipline is to instrument leading indicators from week one and report lagging indicators on a six-month delay.
For the executive perspective on which metrics belong on which dashboard — and how to present AI ROI credibly when the board asks — Executive AI Strategy — Level 5 has the measurement brief in the governance module.
Attribution: the hardest part
The hardest problem in AI ROI is attribution. AI works alongside everything else you changed in the same window — headcount, process redesign, tool refresh, market shifts. Isolating the AI-specific contribution is genuinely difficult, and anyone who claims to have done it perfectly is selling something.
Three approaches that work, ranked by cost:
Workflow-level before/after. Pick three workflows you changed with AI. Measure baseline and post-AI performance on each. Report the gap as “AI-attributable gain” with the caveat that other factors may have contributed.
Controlled comparison. Apply AI to one team and not the other, measure the difference. Strong attribution, often impossible because no two teams are actually comparable.
Counterfactual modelling. Ask domain experts “what would have happened without AI?” Build a structured estimate from their answers. Useful as triangulation, not credible as a primary number.
Report all three where you can, label them clearly, and acknowledge the uncertainty range. A “$2M–$4M AI-attributable gain, central estimate $3M, methodology: workflow before/after on three target workflows” is credible. A “we saved $3M with AI” without the methodology gets challenged the moment a board member asks a sharp question.
The full ROI attribution framework — methodology setup, baseline capture, defending the number in a CFO review — is a deeper module in AI in Business — Level 4.
What NOT to measure (the engagement metrics that lie)
Some metrics look like ROI measures and aren’t. Avoid these as primary metrics:
Login counts. Tells you whether people can log in. Tells you nothing about value.
Prompt volume. Tells you the AI tool is being touched. Tells you nothing about whether the touched work is the work that matters.
“Time saved” surveys. Self-reported time savings are unreliable — people over-report what they want to be true. Use observation or instrumentation.
Vendor-reported productivity multipliers. A vendor’s “average customer sees 30% productivity gain” is a benchmark, not your number. It measures the vendor’s customers on the vendor’s evaluation, not your workflows.
The right primary metrics are workflow-attributable, observed-not-self-reported, and tied to a downstream business outcome. If your primary metric can be gamed by people logging in more or clicking faster, it’s the wrong metric.
For the executive brief on which metrics to put in front of the board and which to keep off the dashboard — Executive AI Strategy — Level 5 covers the credibility discipline.
Sensitivity analysis: the number is a range, not a point
A credible AI ROI number is a range, not a point. “$3M central estimate, $2M–$4M range, methodology: workflow before/after on three target workflows” is honest. “$3M ROI” without a range is not.
The range should reflect methodology uncertainty, counterfactual uncertainty, attribution uncertainty, and scaling uncertainty (the pilot gain doesn’t survive contact with scale).
Reporting a single number is a sign that someone rounded their uncertainty to zero for presentational comfort. Reporting a range is a sign that someone measured carefully and is presenting what they actually know.
The right board report format: “Range $X–$Y, central estimate $Z, methodology W, confidence level V.” Anything less is marketing copy, not measurement.
For the executive playbook on presenting AI ROI to the board in a way that survives scrutiny — format, sensitivity framing, questions to anticipate — Executive AI Strategy — Level 5 has the board-credibility brief.
The 90-day measurement plan
For a new AI rollout, a working measurement system at the end of 90 days has these in place:
Workflow selection done. Three workflows identified, baselines captured (cycle time, error rate, output volume), measurement owner named per workflow.
Leading indicator instrumentation live. Prompt reuse rate, workflow adoption depth, output spot-check scores. Weekly dashboard reviewed by the change lead.
Attribution methodology chosen. Workflow before/after, controlled comparison, or counterfactual modelling — pick the one that fits, document the choice.
Lagging indicator tracking started. Revenue influenced, cost avoided, capacity unlocked. Monthly dashboard reviewed by the executive sponsor.
The 90-day measurement plan is not the 90-day ROI report. ROI reports take six to twelve months of data to be credible. The 90-day milestone is a measurement system you trust enough to read — not a number you trust enough to publish.
For the full 90-day adoption framework — measurement setup, governance, training, coalition work — AI in Business — Level 4 is the complete programme.
The honest answer
The honest answer to “what’s our AI ROI?” depends on whether you’re asking it six weeks in or six quarters in.
At six weeks: you don’t have one. You have activity data, sentiment data, and pilot numbers. None are ROI yet. The right answer to a board question at six weeks: “too early to call, here’s what we’re measuring, here’s when we’ll know.”
At six quarters: you have a credible range, methodology documented, attribution chosen, leading indicators that predicted the lagging outcome. A range that survives contact with a sharp board member.
In between, keep the measurement honest. Don’t inflate the early number to please the board. Don’t suppress it because it’s smaller than the vendor promised. Don’t accept a single number when a range is honest. Don’t accept a range without the methodology.
The vendor’s ROI pitch is a marketing claim. Your AI ROI measurement is a finance function. Different disciplines. The second is harder, slower, and more useful.
Working on AI ROI measurement for a specific function or industry context — and want methodology that fits your situation? Tell us what you’re measuring.