Why I stopped using my coding agent like a tool and started running it like a team.
Part 1 of 3. Series: building an AI agent team to ship software.
For the past few weeks I've worked with agents in a way that's nothing like before.
Before that I used a coding agent the way most people do. Open a session, describe the work, get code back, go back and forth until it runs. For a small utility or a script, that's plenty. I get an idea, type a few lines, the agent answers, I skim it, nod, done.
Then I tried the same approach on a bigger project. And it broke.
The agent's code mostly ran. What broke was my grip on the whole thing. Too many pieces moving at once, tasks sitting at different stages, and no way left for me to tell whether something the agent called "done" actually was.
That's when I changed how I think about it. I stopped seeing the main session as the place where code gets written. I started seeing it as an orchestrator.
The main session doesn't do the work with its own hands anymore. It runs a small team of agents, each with one role. It takes a rough piece of work from me, splits it into tasks, hands each one to the right agent, tracks progress, holds the team to the process, and only commits code after a final acceptance step.
I want to be clear here, because this is easy to mistake for a prompt trick. At heart it's a way of organizing work.
The AI is the same as before. Same model, same tools. What I changed is how the work is arranged around it. And that's the thing that opened up projects I couldn't build the old way.
The reason comes down to one question: who checks the work?
When a single session writes code and also rules that the code is right, nobody checks it. Once the agent says it's finished, I get two options: trust it, or go read every line myself. On a small project I can read it. On a big one I can't. And when I can't, "done" turns into a hollow word.
Splitting the roles fixes exactly that. Every step gets its own gatekeeper. The one who writes the code is separate from the one who decides it's correct.
Let me sit on the breaking point a bit, because once you see it you understand why the roles have to split.
For small work, one agent doing everything makes sense. Few tasks, little state, I hold the whole picture in my head. I know what's done, what isn't, where it's going wrong. If the agent skips a step, I catch it on the spot.
A complex project breaks that in two ways.
The first is volume. This one's being coded, that one's waiting on tests, another has fuzzy requirements, a fourth just failed tests and needs redoing. A single session can't hold that much state straight. It loses track of where a task sits, or it finishes one thing and assumes the whole cluster is finished.
The second way is the painful one, and the real reason I had to change: verification.
When the agent reports "all done," what do I have to go on? No step checks whether what it built matches the plan. Nobody reruns the thing to confirm each criterion. Worse, the same agent that wrote the code also gets to declare the code correct. That's a bare conflict of interest. The student grading their own exam always gives themselves a good mark.
I hit this exact scene more than once. The agent declared a feature complete, I believed it, two days later I opened it up and saw it had missed the requirement from the start. It wasn't being sneaky. It just had no second pair of eyes.
So I decided to build a team. Each agent gets its own role with clear borders. And the main session plays orchestrator: spin up the agents, divide the work, enforce the process, commit only after the final acceptance step.
In my current setup, I split the work across four roles.
Product Manager. Takes a rough piece of work and turns it into something buildable. A spec with a user story, acceptance criteria, test scenarios. This is the role that turns my vague sentence into a clear brief for the whole team. After the code and QA are done, the PM comes back one more time, looks at the result through the user's eyes, and decides whether it's truly finished.
Software Engineer. Writes the code, and writes the tests for that code. This role doesn't get to rule that its own work is right. It just builds, and pushes the result to the next step.
Tester. Runs those tests, checks every acceptance criterion, reports pass or fail with evidence. I lean hard on that word, evidence. QA doesn't get to say a breezy "looks good." It has to show what ran, how it ran, what came out.
On-Call Engineer. Watches CI/CD after the code is pushed, patches things when the pipeline goes red. This is the easiest role to forget, but skip it and code that's "done" on your machine can still break the shared build.
Each role carries one narrow slice of responsibility. Sounds fussy, even slow. But that narrowness is exactly what changes the quality of the output.
Two reasons.
One, narrow slices make skipping a step much harder. One agent doing everything jumps around easily: write code, commit, forget the tests, forget acceptance, because it's in a hurry to reach the finish. When each step has its own role standing there waiting, a skip shows up immediately. A task can't jump from SWE to commit without passing through QA, because QA is a physical link in the chain.
Two, narrow slices make tracing bugs easier. When something's wrong, I know which role to ask. Fuzzy spec, that's the PM. Tests that miss a bug, that's QA. I'm not digging through a fog of "the agent did something for two hours." Each role leaves its own trail, so I trace back to the broken point much faster.
But the gain I value most is this: the role that writes code and the role that grades code are two separate roles. The maker is pulled clean off the grader. That's how "done" gets its meaning back, instead of being the agent patting itself on the head.
This is also where I think a lot of people working with agents are missing the point. They pour their energy into finding a stronger model, a slicker prompt. But my problem was never a weak model. My problem was that nobody verified the work. And verification is an organizational matter, solved by arranging the work around the agent.
I noticed one more thing, and it matters for how I work: this four-role structure isn't tied to any one project.
Once I've defined each role, I carry the whole set over to a new project. The brief changes, the language changes, the technical constraints change, but the four roles and the borders between them stay put. I don't have to think it through from scratch each time. That's the kind of investment I like: spend the effort once, reuse it many times.
The four roles are half the story. The other half is how a task moves through each role in the right order, how to run several tasks in parallel without sitting and watching, and how to track all of it. That's part 2.
If you take one idea from this, take this: the power comes from how you organize around the agent, so that every piece of work has someone to verify it.
1 to 2 emails a month. Easy to unsubscribe.
1 to 2 emails a month. Easy to unsubscribe.
Comments