BlogPractice

What does an AI coding agent actually cost per user story?

4 min read

Nobody can tell you what an AI agent costs per user story, because it depends on your model choice, your story size and how often the agent has to guess. What you can do is make the number measurable: attach every run to the work item it was made for, put four ceilings in place before you need them, and read your own figure after one sprint instead of trusting anyone’s benchmark.

The honest answer to the question in the title is “we do not know, and neither does anyone quoting you a figure”. The cost of an agent finishing a user story depends on which model does the code work, how large your stories are, how good your acceptance criteria are, and how many times the agent has to iterate because something was ambiguous. Those vary by an order of magnitude between teams.

What you can do is stop guessing. Model usage is the first cost in software delivery that varies per work item and can be attributed to it exactly — which is either a reporting nightmare or the best cost visibility the industry has ever had, depending entirely on whether you attach the number to the ticket.

Where the money actually goes

Three things dominate, and only one of them is the one people worry about.

  1. Iteration caused by ambiguity. An agent that has to infer what “handle errors gracefully” means will try, get feedback, and try again. Each loop is a full context window. This is usually the largest single factor and it is a specification problem wearing a cost problem’s clothes.
  2. Context size. Everything the agent is shown is paid for on every call. A context package assembled per role and per budget costs a fraction of “here is the repository”.
  3. Model choice per kind of work. Classifying a comment on a frontier reasoning model costs roughly what writing the feature should. High-volume routine work is where a smaller model saves real money.

Four ceilings, from small to large

CeilingWhat it stops
Per taskA runaway loop on one item. Reached, the agent stops and asks a person how to continue rather than continuing to spend.
Per sprintA whole iteration quietly consuming next quarter’s budget.
Per month, per organisationThe bill as a whole. Counted before the call, not after — a ceiling checked afterwards is a report, not a limit.
Per agent loopIteration and repetition limits, so an agent going in circles stops being expensive as well as useless.
Sprint planning: the sprints as tiles, the current sprint with its items, and the form for the next one.
Budget at planning time. Capacity in points for the people and a budget in euros for the agents — a sprint has to fit inside both, or it stalls halfway with no warning.

Making the number readable

An invoice tells you what AI cost last month. It cannot tell you whether that was worth it, because “worth it” is a per-feature question. The unit that makes the number decision-relevant is the work item: what did this story cost in model usage, next to what it cost in people’s hours.

Progress of a project: what is finished, waiting time on people, AI cost next to hours, and a burndown.
AI cost beside human hours. Kept deliberately apart, because averaging them into one “cost of delivery” figure destroys exactly the comparison you need.

Five ways to bring the bill down

  1. Split work by kind. Routine process work does not need a frontier model. This is usually the biggest single saving available.
  2. Write sharper acceptance criteria. Fewer iterations, and the improvement compounds because tests are derived from the same sentences.
  3. Make stories smaller. If tasks regularly hit the per-task ceiling, that is a decomposition signal showing up on the invoice.
  4. Answer questions faster. Every hour an item sits blocked is context that has to be reassembled when work resumes.
  5. Bring your own key for code work. It moves that spend onto your own provider contract entirely.

What to measure after one sprint

Three numbers, and they are worth more than any published benchmark because they are yours: median AI cost per finished story, the share of items that hit the per-task ceiling, and the ratio of AI cost to people-hours on the same items. The first gives you a planning figure; the second tells you whether your stories are too big; the third tells you whether any of this is paying for itself.

How much does an AI agent cost per user story?

There is no general figure — it varies by model choice, story size and how much the agent has to infer. A reasonable planning assumption to start from is a few tens of eurocents per item, replaced after one sprint by your own measured median.

What is the biggest hidden cost of agentic development?

Iteration caused by ambiguous requirements. An agent that has to guess what “done” means will try repeatedly, and every attempt is a full context window. Better acceptance criteria are the cheapest optimisation available.

Should AI usage be capped per task or per month?

Both, and for different reasons. A per-task cap stops a single runaway loop and turns it into a question for a person; a monthly ceiling protects the budget as a whole. A ceiling that is only checked at invoicing time is a report rather than a limit.