"Is the agent actually working?" is a harder question than it should be — mostly because teams measure the wrong things. Model quality, response speed, and demo-day magic don't pay for anything. Four numbers do. If you track these from day one, you'll know within a month whether your agent is an employee or an expense.

Metric 1: Hours returned

The core unit of agent ROI is time given back to humans. Measure it honestly:

  • Count completed workflow runs, not messages or tokens. One run = the outcome a human used to produce (a review queue prepared, a status report compiled, a ticket resolved).
  • Multiply by the human time that run replaced. Time the manual process before you automate it — not your optimistic memory of it. If preparing the review queue took a person 45 minutes and the agent did 20 of them this week, that's 15 hours returned.
  • Subtract the human time the agent still consumes. Review, approvals, corrections, and handling escalations all count against it. Net hours returned is the real number.

Metric 2: Error rate

An agent that's fast but wrong is a liability with good marketing. Track errors per completed run, split into two classes:

  • Caught errors — the mistake reached a human checkpoint and got stopped. These cost review time but nothing else. A rising caught-error rate means the agent is degrading or the workflow changed.
  • Escaped errors — the mistake made it past review into the world. These have real costs: the wrong invoice chased, the bad reply sent. Escaped errors should be near zero, and every one should trigger a guardrail fix, not just a correction.

Benchmark against the human error rate for the same work — humans make more mistakes on repetitive tasks than anyone likes to admit. A well-built agent with checkpoints should beat the manual baseline on escaped errors within weeks.

Metric 3: Escalation rate

Human-in-the-loop doesn't mean "a human does half the work." Track what fraction of runs require a human beyond a quick approval:

  • Approval-only touches — review and click. Healthy; this is the design working.
  • True escalations — the agent got stuck and a human had to do the step. This is the number to watch. A new agent might escalate 30–40% of runs. A mature one should settle under 10–15% for a well-scoped workflow.

If the escalation rate stops improving, either the workflow's scope is too broad (shrink it) or the process itself isn't as defined as you thought (document it properly first).

Metric 4: Cost per completed workflow

Add up everything — model usage, infrastructure, the build amortized over its expected life, maintenance, and the human review time — and divide by completed runs. Then compare it to the fully-loaded cost of doing one run manually (hours × loaded hourly rate).

For a typical back-office workflow, the agent's cost per run lands at a fraction of the manual cost, and the gap widens as volume grows because the marginal cost of the next run is close to zero. That's the actual ROI story: not "AI is magic," but "marginal cost collapses once the workflow is engineered."

The one-tab ROI model

Build this in a spreadsheet before you deploy, and update it monthly:

  1. Baseline: runs per month × human minutes per run × loaded hourly cost = current monthly cost.
  2. With agent: (runs × agent cost per run) + (approval touches × 2 minutes × loaded cost) + (escalations × human minutes × loaded cost) + monthly maintenance.
  3. Net: baseline minus with-agent, plus a line for quality delta (escaped-error cost, before and after).

If the net isn't clearly positive within the first quarter, the workflow choice or the build is wrong — not the concept.

The metric that doesn't matter

Adoption theater: logins, "engagement," prompts sent. Agents aren't software people log into — they're staff that produce outcomes. If the outcome metrics above are healthy, nobody needs to open a dashboard. If they're not, no dashboard will save you.

Want an ROI model built on your actual numbers?

Bring one workflow and its real volumes. We'll build the model with you and tell you what the agent has to beat.

Request an automation assessment →