Every piece of code encodes what its authors understood about the problem at the time they wrote it. That understanding improves. The team ships more features, sees how data behaves in production, runs into edge cases, and ends up with a better model of the domain. The code does not update itself to match. Six months later the abstractions are slightly off, the module boundaries reflect an older mental model, and some of the names refer to concepts the team has since replaced.
This is the normal state of a codebase, and it is the reason refactoring exists. Refactoring is the work of bringing the structure of the code back in line with the current understanding of the domain. It is a design activity. Treating it as cleanup, or as debt to be repaid in some later sprint, misses what it is for.
This article is about what changed now that coding agents can execute refactorings on their own, and about the part that did not change.
Cosmetic and structural refactoring#
Two different kinds of work both get called refactoring. The first is cosmetic: renaming variables, extracting helper functions, splitting a large file into smaller ones, flattening nested conditionals. It improves readability locally and leaves the design as it was.
The second is structural. It starts from questions about the design itself. Why does this service know about seven domain concepts? Why are the rules for one business decision spread across three layers? Why does a small feature require changes in twelve files? Why does testing one function require mocking half the system? Answering these means rebuilding the domain model with what you know now, and then moving code until its shape matches that model.
Cosmetic refactoring is useful, but it cannot fix a wrong model. A 1,200-line service split into four 300-line services still has the same design, with more indirection between the parts.
Bad structure used to hurt#
Before coding assistants, a wrong structure made itself felt directly. Adding a feature meant threading a parameter through ten functions, passing props through seven components, or editing a dozen files for a single change in behavior. The developer doing that work felt the friction, and the friction carried information: it pointed at the place where the design no longer fit the problem.
Good structure gave the opposite feedback. After a refactoring that put responsibilities in the right places, the next feature often needed a small change in one location, and the rest of the system followed. Less code, fewer files touched. Experiencing that difference first-hand is how most developers learned to care about design in the first place.
The friction moved into the diff#
When an agent implements the feature, the developer no longer does the threading or the prop drilling. The agent does it, following the patterns already present in the codebase. The feature works, the tests pass, and the developer never experiences the cost. Three hundred lines take the agent about as long as thirty.
The structural problem is still there, and it now has three hundred more lines depending on it. Repeat this across months of feature work and the codebase fills with code that expresses the wrong model, and every addition makes the eventual structural change larger.
The signal has not disappeared, though. It moved into the diff. A feature that should have been a small change and instead touched twelve files carries exactly the information a developer used to get from doing the work.
In practice, that diff mostly goes unread. Teams running agents in a loop judge the result by whether the tests pass, and a green thousand-line pull request gives nobody a reason to open it. Reading it would take longer than the agent took to write it, and there will be another one tomorrow. The one place where the structural cost is still visible is the place nobody looks.
The agent will not point it out either. I wrote about this in Sure thing, here's your solution: a model asked for a feature builds the feature. It does not come back and say the structure needs to change first, even when it does.
Execution got cheap#
The bigger change is on the other side. Refactorings that used to take weeks are now cheap to execute. An agent running in a loop against a test suite and a type checker can rename concepts across a codebase, split modules, migrate between APIs, upgrade frameworks, and port entire projects to another language. The automated ports to Rust are the clearest example: the agent iterates until the original test suite passes and the compiler accepts the result.
These projects succeed because the task has an oracle. Correct is defined as behaving exactly like the old code, and the tests and the compiler check that mechanically, so the agent can keep going until every check passes.
The same property explains what such a port does not do. It preserves the design of the source. A tangled codebase becomes a tangled codebase in a different language, because nothing in the task asked for a different structure and nothing could have verified one.
Choosing the target has no oracle#
Structural refactoring has two parts: deciding what the structure should be, and transforming the code into it. Agents have made the second part cheap. The first part has no mechanical check. Tests verify behavior and the type checker verifies consistency, but nothing verifies that the domain concepts are the right ones, that the boundaries sit in the right places, or that the next feature will fit.
That decision was always the hard part of refactoring, and it is now nearly all of it. The person who sets the target and reviews the result determines the quality of the codebase. The agent determines how quickly it gets there.
This also changes the economics behind an old complaint. Refactoring tickets used to lose to feature tickets because they were expensive and their value was hard to argue for. When the transformation costs an afternoon of agent time, the cost argument mostly goes away. What remains is whether anyone has decided what the system should look like.
Who holds the mental model#
There is a second cost to not implementing things yourself, and it is harder to see. Writing the code is how most engineers build a working model of a system: where things live, what depends on what, which parts carry the weight. Choosing a refactoring target depends on that model, and so does judging whether a result is any good. Reading code builds it too, but mostly for people who have written enough code to know what they are looking at.
The usual objection is that this has happened before. We wrote assembly, then C, then JavaScript, and every step handed more of the work to a machine. The comparison fails on one point. A compiler translates deterministically, according to a specification of the language, and the same input produces the same output. The source code remains a precise description of what the program does.
The layers below never stopped mattering, either. In C and C++, the code you write shapes the generated assembly directly: memory layout, aliasing and branch structure all limit what the compiler can do. In JavaScript the influence is less direct but just as real. Add properties to objects of the same kind in a different order and V8 gives them different hidden classes, and a call site that sees too many shapes stops being optimized. Use delete on an object and it can drop into a slow dictionary representation. Put a float into an array of small integers and the array's elements kind changes, and it never changes back. Good code always meant understanding what your code turned into. Each layer was a translation from one precise language into another.
A prompt is not that. A specification in natural language is far harder to make precise than code, and the model filling in the gaps is not deterministic. Hand it the implementation, and the only precise description of the system is code that nobody wrote and few people read.
Reviews without a reference#
Asking a model to check the last few commits against your guidelines, rank the code, or suggest refactorings is useful, and I do it constantly. Deciding whether its answer is right is the problem.
Ask for the same review twice. In my experience the second answer usually revises the first: it has changed its mind about one part, decided another should be different, and suggests a structure the first review never mentioned. Ask a third time and you get a third version. Each answer is plausible, and none of them is anchored to anything, because the model has no fixed picture of what the architecture is supposed to be.
The engineer is supposed to supply that picture. Without a mental model of the intended architecture, there is no basis for choosing between the three reviews, and picking one means letting the model decide. That is not always a disaster. But every decision made that way makes the next one harder to evaluate, because the system drifts further from anything a person holds in their head.
This is also part of why junior engineers are having a harder time. A senior engineer who hands implementation to an agent still draws on years of building systems by hand. A junior engineer who never implemented much has nothing to compare the agent's output against, and the job increasingly consists of exactly that comparison.
Refactor toward the feature#
Fowler's preparatory refactoring still describes the right workflow, and it is cheaper to follow now. When a feature does not fit the current structure, change the structure first so that it does, then build the feature. Where old code is too large to change in one step, put a boundary in front of it, an adapter or a facade, and let new code depend on the boundary while the old code is replaced behind it. Fowler calls this a strangler fig. Every later change to the old code moves it a little closer to the new structure.
The workflow depends on someone deciding that the structure has to change before the feature goes in. An agent will not make that call. It has to come from whoever assigns the work.
With agents the sequence is the same, with one addition: write the target down. An agent asked to clean up a module will do cosmetic work. An agent given the intended boundaries, the concepts that should exist, and what each part is allowed to know will do structural work, and the test suite confirms along the way that behavior was preserved.
Building the judgment#
Deciding targets well is a skill, and the most effective way I know to build it is code review taken seriously. Style comments and suggestions to replace || with ?? do not build it. Asking whether a change belongs where it was put, whether three similar pieces of code point at a missing abstraction, and whether the approach is right at all, does.
This applies to reviewing other people's code, to receiving reviews of your own, and to the agent diffs you do read. Those are a good training ground, because the agent reproduces the existing structure faithfully and shows you, in volume, what that structure costs.
Reading does not scale#
Judgment built in review still has to reach the code, and at agent volume, review is a bottleneck nobody is going to staff. A target architecture written down in a design doc or an AGENTS.md file is a suggestion. The model follows it when it fits the request and drops it when it doesn't, and nothing tells you which one happened.
For the structure to hold, something that does not get tired has to enforce it. Types, lint rules on module boundaries and dependency checks already do this for small parts of a system. The further that goes, the less a person has to read. When what the model writes is constrained enough for a machine to verify, tooling can check most of it and point a reviewer at the few places that still need judgment. What is left for the reviewer is the part only a person can do: deciding whether the structure is right.
