Part of a series on how software architecture shapes AI driven code degradation. This is a progress note on two of ten chains on the review arm.
Two chains is not a result. It is a reading. The full study reports ten runs per cell with confidence intervals, and nothing below is a substitute for that. What two chains can do is tell me whether the arm is behaving, and whether a tidy story I could have told after one chain survives a second. This time it did not, and that is the interesting part.
The review arm, briefly
The fifth condition, reviewed, is the one endogenous signal in the set. The model's own architectural taste, wired up like a pull request. Each checkpoint runs three steps. The author implements the change with a neutral prompt. An independent, read-only reviewer in a fresh session critiques the structure metric-free and cannot run tests or touch the code. Then the original author is resumed with full memory of its own change and the reviewer's notes and it fixes what it agrees with and pushes back on the rest.
The question for this arm is narrow. Left to accumulate changes an AI author lets complexity concentrate. Can a second instance of the same model, looking only at structure, push that back? And does it hold on Spring but not on OfficeFloor the way every earlier lever did?
What two chains show
The two architectures behaved as differently as they do in the paper but the shape of the difference is the headline.
OfficeFloor barely noticed the reviewer and the two chains agree to the decimal.
| OfficeFloor, final phase | chain 0 | chain 1 |
|---|---|---|
| biggest-class WMC | 55 | 55 |
| PMD god classes | 0 | 0 |
| impact added per change (additive) | 15 | 21 |
| cold-read recall | 0.58 | 0.63 |
That is the paper's durability result reproduced under a fifth previously untested lever. The additive architecture holds its shape whether the instruction is nothing, a prompt, a tool, a formula or a reviewer. Architecture not the review is what keeps it flat.
Spring is where the tidy story died. Here are the same two chains.
| Spring, final phase | chain 0 | chain 1 |
|---|---|---|
| biggest-class WMC | 78 | 166 |
| PMD god classes | 1 | 2 |
| impact added per change (additive) | 27 | 265 |
| cold-read recall | 0.68 | 0.61 |
So at two chains the statement is not "AI review tames Spring." It is "AI review can and does not dependably." The mean of the two chains sits below the control, but the spread is enormous and two chains cannot separate it from the control at all. This is precisely the plasticity the paper describes, now visible from the inside. Point a stochastic reviewer at a plastic architecture and you get stochastic placement. Point it at a durable one and you get the same thing every time, because there was never much to move.
None of this bought safety. Spring logged 6 and 4 true regressions across the two chains, no better than the control's roughly 4, and OfficeFloor 2 and 4. Structural stability did not buy correctness in the paper and neither so far does review.
What it cost
The review arm is the most expensive in the study, because every checkpoint is now three model turns instead of one: author, reviewer, and a resumed fix. Per chain of sixty checkpoints.
| per chain | author | review | fix | total |
|---|---|---|---|---|
| Spring reviewed | ~$85 | ~$13 | ~$35 to $43 | ~$135 to $142 |
| OfficeFloor reviewed | ~$93 | ~$16 | ~$47 to $57 | ~$155 to $165 |
The control, author-only, runs around $88 on Spring and $83 on OfficeFloor for the same sixty checkpoints. So the loop adds roughly half again to the bill. It is doing real work for that money: pooled over both chains, the reviewer asked for changes and the author ran a fix on 40 of 120 Spring checkpoints and 52 of 120 OfficeFloor checkpoints, and passed the change on first look the rest of the time. This is not a rubber stamp. It is a negotiation that happens to be costly and, on the evidence so far, structurally unreliable on the architecture that needs it and unnecessary on the one that does not.