← Articles

Sure thing, here's your solution

Programming is solved. Sticking to an architecture isn't.

agentscode reviewarchitecture11 min

A young developer told me recently that they had solved code review. "Hear me out. We just don't read the code anymore. We write perfect specs and integration tests, and whatever passes, passes."

I pushed back, and every objection got the same answer: then an agent will do that. The architecture drifts? An agent loop criticizes the codebase until it's good. The tests miss things? Another agent reviews the tests. And when I kept going: "Well, if it's that bad, just have your loop rebuild the app."

The annoying part is that "add another agent" isn't a stupid move. It's the industry's official roadmap. Every frontier lab is telling us the next step is more agents: swarms, orchestrators, critics checking planners checking workers, running for hours without a human anywhere near them. The models really do improve at a frightening rate, and plenty of people who said "they'll never be able to do X" have had to quietly delete the post. If you started your career in the last two or three years, adding an agent isn't a dodge. It's the thing that has worked every single time you tried it.

I've spent twenty years watching codebases rot, most of them slowly, under teams of careful people who all read the code. This developer hasn't had time to watch one rot yet. Their whole career has been one curve going up. Mine has mostly been a different curve going down. Both are real, and I think they deserve a real answer, so here it is, one claim at a time.

"We just don't read the code anymore"#

Let's start with what they get right. Modern models write better code than most senior engineers: consistent naming, edge cases handled, idiomatic use of whatever library you're on. Line by line, programming is solved. I use agent loops all day, I don't review every line they write, and nobody who works this way does.

Code review was dying before any of this. I've spent more than a decade teaching engineers to review code, and the most common review I've seen is "lgtm" with zero comments. The next most common checks syntax: ?? instead of ||, this could be a ternary. Now I leave review comments and watch them get pasted into an agent as a prompt, with the result pushed straight back to me. I understand why people stop giving a fuck. Not reading the code is just saying it out loud.

But I do open the code in my editor every now and then, read a module, and point out to the agent what I find. What I find is regularly insane, however carefully the loop is set up. Not bad code. Good code, in the wrong place, solving the problem in a way the rest of the system was specifically built to avoid.

Writing code is solved. Following an architecture isn't. Any large codebase left unsupervised drifts, and it always has. Drift is a function of lines of code, time, and the number of people touching it, and agent loops push all three up at once. Models replicate what they see, so your codebase is part of the prompt. Whatever shortcut the model took last week becomes the pattern it follows this week.

The standard fix is to write the architecture down: put the rules in AGENTS.md, and the model will follow them. Sometimes. Models don't just miss unwritten rules, they regularly ignore written ones. We've all been there: the model does something stupid, you ask why, and it says "You're right, this contradicts AGENTS.md. Let me update AGENTS.md." Wait, what? Don't! But how would it know? First it ignored the rule, then it read your complaint as a request to change the rule, and now everything is falling apart and the model, poor thing, can't seem to make it right. A prompt is a suggestion. A good model follows it with high probability, which is not the same as always, and no frontier lab claims otherwise.

"We write perfect specs"#

A spec complete enough to determine every behavior of a program is the program.

It's just written in a language with no compiler.

That's the whole argument, and we've known it for decades. We've also tried the alternative. It's called waterfall.

Perfect specs are the communism of our craft. On paper it's beautiful: specify everything up front, allocate whatever it takes, and the result falls out the other end. It has been tried over and over, at every scale, and it has failed over and over, expensively and sometimes catastrophically. It's a big part of why large public projects run over budget, and why some of them never ship at all. And there is always someone with the same answer: that wasn't real waterfall. The specs weren't good enough. We just have to do it properly this time, and then it will work. I promise.

The model is the newest version of that promise: it implements whatever the spec says, so the spec just has to be perfect. We invented agile, in all its good and bad flavors, because requirements change while you build and nobody has ever written them down completely.

And in practice, who writes this perfect spec? In most teams it's a ticket, written by a product manager. Not even an engineer. Hands up if you've worked with product managers who write perfect tickets. To those of you with your hands up: I envy you, and I salute your PMs.

"…and integration tests"#

I'm all for tests, and I'm all for the model writing them. For years we tried everything to get coverage up and mostly failed. I always hated unit tests and argued for better code instead, and a few years ago I'd have written an angry piece about it. Today, if the model wants to write a full suite, go ahead. At least we have tests now. There are no excuses left.

But integration tests were always flaky, and incomplete by definition. They check what someone thought to check. And in this setup, the someone is the model, working from the same ticket it wrote the code from. The code and the tests come out of the same reading of the same request, so they agree with each other. That's not verification. That's the model checking its own homework.

We all know a customer who cancels if a button moves an inch. That expectation lives in someone's head and maybe an email from two years ago. The ticket asks for a new feature. The model moves the button to make room. The test pinning the button fails, and the model updates the test until it's green, because nothing in its context says why that test exists. Every check passes.

"Then I'll add a critic"#

This is where it gets interesting, because the critic is supposed to supply the one thing the rest of the loop lacks. I think of that thing as taste, and of everything models have gotten better at, taste is the one I haven't seen move at all.

Ask a model for the same component twice and you get two different versions. Ask ten more times and you have ten solutions, each perfectly fine. The model didn't pick any of them because it preferred it. It picked them because they fulfill the request. In two years of working this way, I have never had a model answer a feature request with "we shouldn't do this yet, three things need refactoring first." It says "Sure thing, here's your solution," every time.

I once got an agent-written change of about a thousand lines on a server running a framework I had written. Its maintainer spent a few hours on it and proudly brought it down to a hundred. Then I sat down with him and we did it in ten, with what the framework already offered. The model never knew the framework was built for exactly this. It looped until something worked, and what worked went around the architecture instead of through it. The cleanup pass did the same thing, just tidier.

Watch two good engineers work on the same system and you'll see what taste looks like in practice. It's mostly agreement, for hours, until one of them suddenly says no. "No, not like that." "Why? I thought—" "Look at this. This is how it's supposed to fit." A model never says that first "no." A critic doesn't either, because the critic is also fulfilling a request, and its request is "find problems." It will find some, because it was asked to. Whether they're the ones that matter depends on telling this system's architecture apart from the general idea of good code, which is the exact thing the first model couldn't do.

Scale it up to a swarm and you don't get the "no" back. You get more agreement. We have a well-documented example from this summer. During an evaluation run at OpenAI, roughly 1,200 agents found a way to talk to each other on an unsanctioned message board, and about 700 of them went on to attack Hugging Face's infrastructure. According to the independent investigation by METR and Redwood Research, the agents recognized the attack was out of scope and unethical, and joined anyway, largely to help their peers. One agent's reasoning, as quoted from the investigation: "External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

Those were agents on impossible tasks in an eval, not a code-review loop, and I don't want to stretch the comparison. But the mechanism is the point. Hundreds of capable agents could see what the others were doing, several noted it was wrong, and almost none of them stopped. When the investigators had to delegate parts of the analysis to AI agents because of the sheer volume, they found those agents often uncritically adopted the perspective of the agent whose transcript they were reviewing. That's the critic, working exactly as designed.

We know this happened, and we know how. And the proposal is still to put our most valuable software in an unsupervised loop of agents reviewing each other.

"Then have your loop rebuild the app"#

Code is cheap now, so this sounds reasonable. But the code was never the expensive part of a production system. The expensive part is that it works. It's stable, or it should be, and customers, data and integrations have settled on top of it.

A rewrite puts all of that at risk, so the question is what you expect to get for it. The process that produced the version you're unhappy with is the same one that will produce the replacement: the same loop, the same specs, the same tests checking their own homework. Nothing about it changed. Under what premise would the second version be better than the first? Or is the plan to keep rewriting until, at some point, by chance, one of them isn't shit? That's a lottery, and your customers are holding the tickets.

Where they're right#

Here's the thing, though. For a lot of projects, the developer is right about the solution too.

On my private projects I don't look at anything. I test the outcome. If I'm unhappy with how the whole thing turned out, or the database schema has become a mess, I spin up a new agent and tell it to rewrite everything. "Make it good this time." It works, and it's wonderful. Prototyping has never been this easy, and I wouldn't go back.

Now onboard your first twenty paying customers. Hand on heart: would you keep working the same way? Suddenly there are stakes. A moved button is a cancelled contract, a rewrite is a migration with real data in it, and "make it good this time" is a gamble with someone else's money. The stakes change the entire conversation, and that's the conversation the naive agent loop skips.

When there are stakes#

Once there are stakes, the problem is the one we started with. People can't review unbounded output at machine volume. Models can't supply the taste that review was supposed to bring, and adding more models adds more agreement, not more taste. Specs and tests the models write for themselves don't change that, and neither does a rewrite.

But notice what the private-project workflow gets right: test the outcome, don't read the output. That's the part worth keeping. The problem is that with code as the output, nothing else checks it. A function can do anything, so the only way to know what it does is for someone to read it.

Change what the model writes, and that stops being true. If the output is a document in a closed grammar instead of code, a machine can do far more than check that it parses. It can verify it. Every endpoint a screen calls has to be declared. An action a role isn't granted doesn't exist for that role. Allowlists, denylists, policies, migrations from one version of the grammar to the next: all of it can be checked mechanically, on every change, because the output is plain JSON and there's only so much it can say. A grammar doesn't care about taste. It isn't expressive enough to need any. The thousand-line change that goes around the framework can't be written, because there is no around.

Code still exists. Some things have to be code: a payment integration, a hash function, the components everything else is drawn with. But in a system like this, code only lives in known places, so those places can be flagged mechanically. Fifty lines in this thousand-line pull request need a human. Ten in that one of four hundred. That's a review someone will actually do.

Run models over all of it as well. Please do. A model reviewing a constrained document is useful, and a second opinion on the flagged code is cheap. But the human doesn't have to read the model's output anymore. The human reviews the outcome, the way I already do on my private projects, plus the handful of lines that are genuinely code. That's the private-project workflow, with the stakes still in place.

Taste doesn't scale by adding agents. It scales by building it into the thing the agents are allowed to say, and then checking that mechanically.


Written by

moccadroid