BlogPractice

Vibe coding works alone and breaks in a team — here is exactly where

4 min read

Vibe coding — building software by describing it to a model and steering the result — is real, fast and undersold by its critics. It breaks on three specific things: memory that ends with the session, agreements that exist in one person’s chat history, and a definition of done that is whatever the model said. Each has a concrete replacement, and none of them require giving up the conversation.

Let us start by conceding the point, because most critiques of vibe coding do not. Describing what you want, watching working software appear, and steering it by conversation is a genuinely faster way to build certain things. Prototypes, internal tools, the first version of anything: it works, and pretending otherwise makes the rest of the argument sound defensive.

What it does not survive is other people. Not because the code gets worse but because three things that a solo builder holds in their head have nowhere to live once a second person shows up.

Break 1: the memory ends when the session does

Everything you decided is in a chat log, and a chat log is a terrible database. You cannot query it for “what did we agree about password lockout”. You cannot diff it. You cannot tell which of two contradictory statements was the later one. And nobody else can read it, because it is full of half-thoughts you never intended anyone to see.

What replaces it: work items and tickets. Not because ceremony is good but because a decision needs an address — something you can link to, comment on, and find in a year. The conversation stays; it just stops being the only place the outcome exists.

Work item US-104 with its acceptance criteria, Definition of Done, and the conversation about this ticket.
A decision with an address. Acceptance criteria, Definition of Done and the conversation about this specific item — findable without scrolling anyone’s history.

Break 2: two people build two versions of the truth

You and a colleague each have a session. Each session has its own context, its own assumptions, and its own idea of what “the user model” means. Nothing tells either of you that the other has diverged until the merge, and by then both branches are internally consistent and mutually incompatible.

This is not a tooling gap that a better assistant closes. A session is private by design. The fix has to be something both sessions read from and write to — which is the definition of a board.

Break 3: “done” is whatever the model said

A model reporting that it implemented and tested something is a claim about its own work. Solo, you catch it: you run the thing, you notice. In a team the claim propagates — into a stand-up, into a status report, into a release note — and the moment where someone would have caught it never arrives, because everyone assumes someone else did.

What replaces it: conditions computed outside the model. CI reported green. A pull request is open. A review was submitted by an actor that did not write the code. A deployment to TEST happened. A human tested it. Each is a fact with a source, and none of them can be produced by the party being checked.

What you keep

All of the speed, if the transition is done properly. You still describe what you want in a conversation. An agent still writes the code. What changes is that the instruction lands as a work item, the decision lands on a ticket, and the finish is checked by something other than the thing that did the work.

  • Keep: describing intent in plain language instead of writing tickets by hand.
  • Keep: letting an agent do the first draft of a decomposition, a description, a test case.
  • Add: an address for every decision, so the second person can find it.
  • Add: conditions on the transitions that matter, evaluated outside the model.
  • Add: a cost per work item, because usage-priced work needs attribution to be manageable.

The honest cost of the transition

It is slower on day one. Writing acceptance criteria takes longer than saying “make login work”. The payoff shows up at the first handover, the first regression, the first customer question about why something behaves the way it does — and those arrive sooner than people expect, usually in week three.

Is vibe coding bad?

No — it is fast and effective for one person on one project, especially early. It breaks on team scale for three specific reasons: session-bound memory, private divergence between people, and a definition of done produced by the party doing the work.

When should a team move away from pure conversation?

When a second person joins, when a customer starts asking why something works the way it does, or when you first cannot answer “was this reviewed and tested” without scrolling. In practice that is week two or three of real use.

Does adding process slow the AI down?

It slows the first hour and speeds up everything after. Agents work better with explicit acceptance criteria than without them — an agent that has a specification does not iterate, and iteration is what actually costs money.