Human in the loop: which decisions must stay with people when agents do the work
“Human in the loop” degrades into rubber-stamping unless you are specific about which decisions stay human and why. Three questions decide it: how reversible is the action, who is accountable if it is wrong, and could the agent have manufactured the evidence for it. Everything that fails those goes to a person; everything else should not, because approvals that mean nothing teach people to click.
“Human in the loop” is one of those phrases everyone agrees with and nobody implements the same way. In practice it collapses into one of two failure modes. Either every action needs an approval — which produces a queue of clicks that people learn to clear without reading — or the approvals are on whatever the tool happened to make configurable.
Neither is a design. A design says which decisions stay human, and can explain why for each one.
Three questions that decide it
1. How reversible is it?
Reversibility is the strongest single predictor of whether something needs a person. Opening a pull request is reversible — close it. Merging to main is expensive to undo. Deploying to production touches people outside your organisation. Permanently deleting an item destroys the record of why anything happened. The further along that scale, the less negotiable the human is.
2. Who is accountable if it is wrong?
Some decisions are commitments to other people: a sprint is a promise to a customer, an acceptance is a statement that the requirement is met, a release is a claim to users. An agent cannot be accountable for a promise, which means somebody has to be — and that somebody should be the one who made it, not the one who discovers it later.
3. Could the agent have produced the evidence itself?
This is the question people forget, and it is the one that separates a real check from a decorative one. If the evidence for “this is fine” was generated by the actor asking for approval, the approval is only as good as the reviewer’s scepticism at 5pm on a Friday. Evidence with an independent source — a pipeline result, a review by a different actor, a human test — is what makes an approval mean something.
A workable division
| Decision | Who | Why |
|---|---|---|
| Decompose a document into a backlog | Agent proposes | Fully reversible; editing a draft is cheap. |
| Refine a story, write criteria | Agent drafts, person edits | Reversible, but it is the specification everything downstream reads. |
| Plan into a sprint | Person approves | A commitment to other people. An agent may propose. |
| Write code, open a pull request | Agent | Reversible, reviewed, and nothing has shipped. |
| Approve a review | A different actor than the author | Independence is the whole value; self-approval is not review. |
| Merge to main | Person, always | Expensive to undo and the first irreversible step. |
| Release to production | Person, always | It reaches users. No configuration should be able to change this. |
| Accept a delivery | Person, always | A statement about whether the requirement is met. |
| Change the workflow or the gates | Person, always | An actor that can widen its own constraints has none. |
Asking is a first-class action
The most useful thing an agent can do when it does not know something is ask, visibly. In Turnado a question posted on a ticket puts the item on Blocked and an answer takes it off — so the uncertainty becomes a state of the work rather than a silent guess or a stalled process.
That also makes the human bottleneck measurable. “Waiting on a person” is a number on the dashboard: hours per work item that stood still waiting for someone. Most teams find, once they can see it, that their constraint is not agent throughput at all.

Where the boundary should live
If the answer to “can an agent merge to main here?” is “no, because we configured it that way”, then it is yes on the day someone reconfigures it under deadline pressure. Boundaries that matter belong in code, applied after roles are resolved, guarded by tests that fail if the list shrinks.
Which decisions should always stay with a human?
The irreversible ones and the ones that are commitments to other people: merging to main, releasing to production, accepting a delivery, approving a sprint, and changing the workflow or the gates themselves.
Is human-in-the-loop enough to keep agents safe?
Only if the human has independent evidence. An approval based on the agent’s own report is a formality. Pair every human checkpoint with a fact from a different source — a pipeline result, a review by another actor, a test a person ran.
How do you avoid approval fatigue?
Have fewer approvals and make each one carry evidence. Reversible steps do not need one. If you cannot say what a reviewer should check and what they have to check it with, the step is ceremony and it is training people to click.