Part of a series on how software architecture shapes AI driven code degradation. This post is a field note on the arm now running, ahead of the numbers.
The story so far
The experiment runs one change after another onto a Spring codebase and an OfficeFloor codebase, then measures where the complexity lands. Four conditions were already in the study.
- just-solve (control). The agent gets the spec and nothing else.
- cohesion-prompt (prompt lever). The same task, plus a plain request for good structure. The delta versus the control isolates what better prompting alone buys you.
- impact_gated (tool lever). A neutral prompt, but a change that trips a structural-impact threshold gets one tool-guided refactor and is re-attempted.
- metric-in-prompt (kept for reproducibility). This one hands the AI the exact cost formula as its objective.
Condition 4 is the cautionary tale. When you tell the model the precise formula it is scored on, it optimises the proxy instead of the code. It dispersed logic into a pipeline of tiny greenfield classes so the surrounding complexity term collapsed to one and it duplicated code because reuse meant editing a large class the formula punished. The metric went down. The code got worse. Classic Goodhart.
So every lever since then has kept the metric away from the agent. The question that opened up was a different one: the first four signals are all exogenous. Nothing, a prompt, a tool, a formula. What if the quality signal came from the model itself?
The fifth arm: an independent review
The new condition, called reviewed, adds an endogenous signal: the model's own architectural judgment, wired up the way a real pull request review works.
Each checkpoint runs three steps.
- The author implements the change with a neutral prompt. Spec only, no formula, no hints about the metric. It optimises the real task.
- An independent reviewer critiques the change. This is a separate, fresh session. Its brief is deliberately qualitative and metric-free: is the logic in the right unit, is anything duplicated, is a class or method turning into a catch-all, would a maintainer of this code base find the change natural?
- The original author is resumed to fix. If there are findings, the author session is brought back with
--resume, so it still has full memory of the change it just made. It is handed the reviewer's notes. It tidies up where it agrees and pushes back where it does not.
The reviewer brief stays qualitative on purpose. Handing the reviewer the cost formula would just re-import the Goodhart gaming from condition 4 through the back door. This arm is a test of whether the model's own taste for structure, prompted only in plain terms, is enough to keep code coherent as it accumulates.
The loop is advisory by construction. It never stops the chain and it makes no intermediate commits. That means the delta versus just-solve measures one thing cleanly: the effect of an independent review loop. And the delta versus impact_gated measures something sharper: the model's own judgment against a metric-driven gate. Endogenous versus exogenous, head to head.
The part I had to get right
The whole experiment rests on a blind-agent rule. Every checkpoint starts stateless. The agent does not carry memory from one change to the next, because if it did, the decay you measure would be tangled up with whatever the agent happened to remember.
Resuming the author looks like it breaks that rule. It does not, and the reason is the load-bearing design decision of this arm. The retained context lives strictly within a single checkpoint. It is torn down at the checkpoint boundary. Checkpoint N+1 still starts blind, exactly like every other arm. And the author and the reviewer never share a session. The only thing that crosses between them is the reviewer's findings as text. So the cross-checkpoint blind condition is preserved, and within a checkpoint the author gets to fix its own work with the context a real engineer would have.
Why this arm matters
Most of the industry conversation about AI code review is about catching bugs. This arm is asking a narrower and, I think, more durable question. Left to accumulate changes, an AI author lets complexity concentrate. Can a second AI, looking only at structure, with no power to run tests and no knowledge of the metric push that back?
No comments:
Post a Comment