Measuring AI autonomy: five metrics that tell you whether your agents actually help
Most AI productivity reporting measures activity: tasks completed, code generated, suggestions accepted. None of those tell you whether delivery improved. Five metrics do — autonomy rate, waiting time on people, cost per finished story, rework rate, and time from analysis to first tested increment — and each has a specific failure it is designed to expose.
The reporting problem with AI in delivery is not that the numbers are hard to get. It is that the easy numbers measure activity. Tasks completed, suggestions accepted, lines generated — all of these go up when an agent is busy, including when it is busy producing work that a person then has to undo.
Five metrics survive contact with reality. Each exists to expose a specific way an agentic setup can look healthy while being useless.
1. Autonomy rate
What: the share of work items an agent worked on that reached Done without a person having to step in beyond the required approvals. Exposes: agents that technically do work but generate a correction for every action.
Watch it against a baseline rather than a target. A rate of 100% on two items means nothing; a rate that drifts down as story complexity rises is telling you exactly where the ceiling of your current setup is.
2. Waiting time on people
What: average hours a work item stood still waiting for a human — an answer, a review, an approval. Exposes: the actual bottleneck, which in almost every agentic team is not the agents.
This is the metric that changes management decisions most often. Teams arrive convinced they need more agent capacity and discover their items spend most of their life waiting for a review that takes four minutes to do.

3. Cost per finished story
What: median model spend on items that actually reached Done. Exposes: spend on work that was abandoned, and the true price of your model configuration.
Median rather than mean, and finished rather than all: the mean is dominated by the one item that went into a loop, and including abandoned work hides the fact that you paid for it.
4. Rework rate
What: the share of items that come back — reopened after Done, or a bug linked to a story finished in the last sprint. Exposes: the single most common way agentic delivery flatters itself, which is completing quickly and correcting later.
If autonomy rate goes up and rework goes up with it, nothing improved: the work moved from “doing” to “redoing”, and the second one is more expensive because it carries the cost of having believed it was finished.
5. Time from analysis to first tested increment
What: days from a document arriving to something a person has actually tested. Exposes: whether any of the speed reached the customer, or whether it all accumulated in a backlog.
This is the one number a customer recognises. Everything else is internal.
Measuring honestly
- Show “nothing measured yet” rather than zero. A zero claims something; an empty state admits something. The difference matters when a number is going into a board report.
- Keep human hours and agent hours apart. Averaging them into one delivery cost destroys the only comparison anyone actually wants.
- Count estimates that were far off in both directions. Systematically over-estimating is a planning problem too, and it is invisible if you only track overruns.
- Attribute to the work item. Any metric that only exists at project level cannot answer “was this feature worth it”.
What is the best single metric for AI in software delivery?
If you can only have one: median cost per finished story, tracked next to rework rate. Alone, cost rewards abandoning hard work; alone, rework rewards doing nothing. Together they are hard to game.
How do you measure whether agents actually save time?
Compare time from analysis to first tested increment before and after, and check rework rate did not rise. Throughput improvements that come with more rework are not savings; they are deferred costs with interest.
Why is waiting time on people so important?
Because in most agentic teams it is the constraint. Agents produce work faster than people review it, so total throughput is set by the review and answer queue — which means adding agents changes nothing until that queue is addressed.