Three real runs with the agent team, including the one I botched — and what keeps it from drifting.
Part 3 of 3. Series: building an AI agent team to ship software.
The first two parts were theory and process. Part 1 covered why I split the main session into an orchestrator running four roles. Part 2 covered how a task moves through the team and how to keep the loop from dying overnight. This part runs the whole thing for real, across a few kinds of projects, and then tells the story of the time I botched it.
I'm deliberately not painting it rosy. This approach wins on some kinds of work and stumbles on others. Knowing where it stumbles matters more than knowing where it wins.
The work this fits best is a project built from zero.
When there's no old code to respect and no constraints to honor, the agent has the most room. I describe the thing I want, the orchestrator tears it into a stack of tasks, pushes them into the backlog, and lets the team grind through the night the way part 2 described.
Morning comes and most of the backlog has gone the full loop: PM groomed, SWE coded, QA checked, PM accepted. I didn't sit and watch the clock. I just open up and read each task's result, and because every task has its own record with QA evidence, I know right away which ones I can trust and which ones I should look at hard.
This is where the self-running loop pays off clearest. A new project has many independent tasks with few cross dependencies, so two tasks running in parallel rarely step on each other. The more loose, separate work in the backlog, the more this approach gains.
I still have to say one thing straight, so nobody reads this and pictures a free code printer. The quality of the output equals the quality of the spec the PM wrote up front. A loose brief means the agent runs all night and produces a pile of things that work but miss the point. It took me a fair few rounds to learn that the effort poured into the PM step up front comes back several times over at QA and final acceptance.
The second fit is just as strong: small jobs, clear scope, one purpose.
A tracker running serverless. A utility library. A set of benchmark scripts to measure performance. Work like this has sharp borders, easy to describe as acceptance criteria QA can check all the way through.
For this kind I usually don't need many parallel lanes. The backlog is short, so running it sequentially finishes fast too. What I gain most here is peace of mind: with a separate QA and a final PM step, I don't have to reread every line myself to trust that it's right.
What do those two kinds share? Both give the agent a clean space to work in. Little legacy, few hidden constraints, a brief that fits in a box.
Now the part fewer people tell.
I took the same setup into something completely different: rewriting part of an existing system, porting it to a different build. Building on the old base, with a heap of hidden behavior the current code carries.
I botched it. And I botched it exactly where I thought I already understood things.
I rushed and skipped the PM step. I figured "I know this work cold, why write a long spec," so I described it roughly and turned the team loose. Backlog full, loop self-rolling, two tasks in parallel, just like the nights before.
Morning came and I got a pile of code that ran, tests all green, QA reporting pass. Every indicator looked great. But it ported the wrong thing. It kept some behaviors that should've been dropped, and dropped some that should've been kept, because I never wrote out clearly which was which.
The tests couldn't save me, because tests only check "does the code do what we told it to." I told it wrong from the start. QA couldn't save me either, because QA checks against criteria, and the criteria grew out of my loose spec. The whole line ran flawlessly on a broken brief.
The lesson was expensive but tidy: the more a job carries legacy and hidden constraints, the less you're allowed to cut the PM step. The very step I thought was redundant turned out to be the one keeping the whole job on the rails. I pulled off the protective layer with my own hands, then blamed the machine.
Looking back, the reason is clear enough.
A greenfield project hands the agent a blank sheet. A rewrite on an old base is the opposite: the hardest part lives in things that aren't written in the code, that live in the head of the person who's lived with the system. Why this spot does it this way. That quirk exists to humor a rare failure case. That constraint is there because some far-off system depends on it.
The agent can't read those things off the code. I'm the one who knows them. And I was too lazy to feed them into the spec. So the agent guessed, and it guessed wrong in a way that looked a lot like right.
This agent team doesn't generate understanding of an old system on its own. It only amplifies the understanding I feed in. Feed it fully and it multiplies that. Feed it short and it multiplies the shortfall, fast and tidy enough that I don't think to doubt it.
Pulling together the nights that worked and the one that didn't, I'm left with three things. Miss one of the three and the agent starts to drift.
One, a clear spec. This is the thing I underrated and paid for. A spec sounds like paperwork, but it's where I pour the understanding the agent has no way to get on its own. A loose spec means every later step runs cleanly on a wrong foundation.
Two, assigned roles. The one who writes the code isn't the one who grades the code. Separating roles is what gives "done" its meaning, and what lets QA send a task back without ego. Part 1 covered this closely.
Three, a fixed process. Every task runs the full sequence, no exceptions. Write the process to a file so the next project carries it over without retraining. Part 2 was about this.
With all three, the agent runs nearly straight and I barely have to watch. Miss the spec and it guesses. Miss the roles and it grades its own work. Miss the process and it skips steps and calls things done too early. The night I botched it was the night I pulled out the first one myself.
If you ask me what actually changed over these weeks, I won't point at a model or a prompt.
What changed is that I stopped seeing the agent as a machine I command and started seeing it as a team I have to organize. Same model, same tools. What I rearranged was the work around it: who builds, who guards, who grades, which step it goes through, what "done" means.
And the overnight automation trick, much as I love it, is just the top layer. It's only safe when the three things underneath are tight. Automation without verification only helps you make wrong things faster, exactly like the night I ported it wrong.
If you carry one line from this whole series, carry this: however strong the agent gets, it only runs straight when you build enough structure around it that every piece of work has someone to check it, and every "done" has evidence behind it.
1 to 2 emails a month. Easy to unsubscribe.
1 to 2 emails a month. Easy to unsubscribe.
Comments