FDE Műhely

For leaders

How to measure AI ROI so finance accepts it

Three groups of metrics matter: cost savings, risk reduction and revenue impact. How to turn each into defensible numbers without cheating.

An AI project survives a budget round if there’s a number behind it. Not “we became more efficient” — currency.

Three groups, each measured differently.

1. Cost savings

The easiest to measure and the easiest to overstate.

The correct formula:

saving = (baseline time − new time) × case volume × hourly cost
       − the system's running cost
       − the new time spent on review

The two subtractions are what people routinely omit.

Baseline time. Measure it, don’t estimate it. That’s why we capture timing data during discovery — without a baseline every later claim is arguable.

New time. Not zero. Exception handling and approval remain. If you automate 85% of a process, real time saved is typically 60–70%, not 85%.

Running cost. Model usage, infrastructure, maintenance. The last one gets forgotten: a live system needs a few engineering hours a month even when it’s behaving.

Careful with “freed-up time”. If freed time converts into neither headcount reduction nor other value-creating work, it’s spare capacity, not a saving. Name it separately, or finance will — rightly — strike it out.

2. Risk reduction

The one people least like to quantify, and often the larger number.

Error rate. How many faulty items got through before, and how many now? This needs a baseline measurement — typically reviewing a retrospective sample. A faulty accounting entry has a known cost: time, possibly late-payment interest, in the worse case a correction filing.

Delay. How many items missed a deadline? If penalties or interest attach, that’s direct currency.

Compliance. If a process becomes logged and searchable, that’s measurable time at audit. Ask your auditor.

Key-person dependency. If one person is the only one who can run a process, that carries risk value. If the knowledge moves into a system and a documented process, it drops. Hard to convert to currency, but it belongs in the risk register.

3. Revenue impact

The hardest, because causation is rarely clean.

What you can honestly claim:

  • Turnaround time. If quoting drops from three days to four hours, and you can show faster responses correlate with a higher win rate, that’s revenue.
  • Capacity. If the same team issues twice as many quotes and there’s demand for them, that’s revenue.
  • Lost opportunities. If you previously didn’t respond to enquiries for lack of capacity and now you do, that’s measurable.

What you shouldn’t claim: that a revenue increase came from AI when other things changed too. Give finance one indefensible number and they will never believe the defensible ones either.

The report that works

Quarterly, one page:

Process: supplier invoice processing
Period: 2026 Q2 · baseline: 2025 Q4

Volume                        4,180 cases     (+8% vs baseline)
Completed automatically       3,512  (84%)
With human approval             521  (12%)
Sent to exception queue         147   (4%)

Average handling time         11 min → 2.4 min
Hours saved                   603
Net saving                    ~€18,000     (running cost deducted)

Faulty items                  2.1% → 0.4%
Items past deadline           86 → 11

Running cost
  model usage                 €1,030
  infrastructure              €  240
  maintenance (6 eng. hours)  €  450

What makes this credible: the cost is in it, and so is what didn’t work. The 4% exception queue is there. A report where everything is perfect is suspicious — and your CFO knows that.

Why this is worth taking seriously

Not for the sake of reporting. Because your internal sponsor will use it when their own review comes around. Give them a page of defensible numbers and they do well too, and the next project starts more easily.

Without it, the project is “some AI thing we did last year”, and there is no next one.


Related: the ROI matrix and token maxing.

Questions on this topic

What should the baseline be?

Measured data from the three months before go-live, not an estimate. That's why we capture timings during discovery: without a baseline, every later claim is arguable.

Can we count freed-up time as a saving?

Only if something actually happens with it. If the colleague moves to other value-creating work, that's real. If the day just gets easier, that's spare capacity — name it separately, because finance won't accept it as a saving.

What does this look like at your company?

If this problem sounds familiar, let's start with one process. Tell us which department burns the most manual hours — we'll come back with a concrete proposal.

Related articles