BlogPractice

The eighteen conditions an AI agent should never be able to skip

4 min read

Guardrails for autonomous agents are usually described as prompts and policies. Neither survives contact with an agent that reads untrusted input. The alternative is boring and effective: conditions that are pure functions of the work item’s state, evaluated outside the model, on every transition that matters — plus a set of permissions that are unreachable regardless of role.

Ask how an agent platform keeps agents in line and you generally get one of two answers: a carefully written system prompt, or a policy document. Both are worth having and neither is a control. A prompt is a request to a system that is designed to be persuaded, and an agent that reads a customer document, a pull request comment or a web page is reading text an attacker can write.

The alternative is unglamorous: put the check outside the model, make it a function of state, and give the agent no way to reach the thing it would need to change.

What a good condition looks like

  • Pure. No model, no network, no clock. The same item state always yields the same answer.
  • Evidence-based. It reads a fact with a source — the pipeline reported, a review was submitted, a deployment happened.
  • Non-self-serving. The party being checked cannot produce the evidence.
  • Explainable. When it refuses, it says which condition is open, in words, on the ticket.

That last one matters more than it sounds. A greyed-out button teaches nothing and gets worked around. A refusal that names the missing condition turns the guardrail into documentation.

Work item US-104 with its acceptance criteria, Definition of Done, and the conversation about this ticket.
Refusal with a reason. Every transition that exists from this status, and word for word which condition is still open.

The eighteen, grouped by what they protect

GroupConditionsWhat it protects against
Specificationacceptance_criteria_present, estimate_present, dependencies_resolvedAn agent starting work on a story that does not say what “working” means.
Commitmentsprint_assigned, assignee_presentWork that nobody agreed to and nobody owns quietly entering the flow.
Uncertaintyquestion_posted, question_answeredAn agent guessing instead of asking — and the wait being invisible when it does ask.
Codepull_request_open, commit_linked, ci_green, review_approvedCode that was never seen, never built, or approved by whoever wrote it.
Testdeployed_to_test, test_instructions_posted, tester_assigned, test_passed“Tested” meaning the model ran its own unit tests and nothing else.
Closuredefinition_of_done_complete, bug_linked, human_review_approvedDone meaning done to the agent rather than done to the team.

Conditions are not enough on their own

A condition constrains a transition. It does nothing about the actions that bypass transitions entirely — merging a branch, deploying to production, deleting the item, or editing the workflow so the condition no longer applies. Those need a second mechanism: permissions that are removed after roles are resolved, in one place, for every agent, regardless of configuration.

Making inaction visible

Guardrails create a new failure mode, and it is worth naming: an agent that is blocked and silent looks exactly like an agent that is working. That is why the queue has to say not only what can be picked up but why everything else cannot — waiting on a human, WIP limit reached, budget ceiling, a status agents do not act on.

The queue: what can be picked up now, the slots taken by AI colleagues, and what stays behind by cause.
Everything that stays behind, grouped by cause. An agent that quietly does nothing is the failure mode of the whole category; naming the reason is the fix.

A minimum set, if you are starting from nothing

  1. Acceptance criteria before work starts.
  2. A pull request and green CI before review.
  3. Review by an actor that did not write the code.
  4. A human test before Done.
  5. Merge and production release unreachable for agents, permanently.

Five rules is a weekend of work and removes most of the ways an autonomous system embarrasses you. The other thirteen conditions are refinements of the same idea.

Why not just write a stricter system prompt?

Because a prompt is an instruction to a system built to be influenced by instructions, and agents read text that other people wrote — documents, comments, web pages. A permission check is a refusal that does not depend on the model’s cooperation.

Does this stop an agent from being useful?

It stops an agent from being the last word. It can still analyse, decompose, write code, open a pull request, review someone else’s code, derive test cases and draft documentation. What it cannot do is certify its own work or take the two irreversible steps: merging and releasing to production.

What is the most commonly missing guardrail?

Separation between the actor that writes code and the actor that approves it. Teams often run one general-purpose agent for both, which turns the review gate into a rubber stamp while leaving every dashboard looking healthy.