Monday, 28 September 2026

Can an AI review keep AI code clean? Adding a review loop to PetClinic-Evolve

Part of a series on how software architecture shapes AI driven code degradation. This post is a field note on the arm now running, ahead of the numbers.

The story so far

The experiment runs one change after another onto a Spring codebase and an OfficeFloor codebase, then measures where the complexity lands. Four conditions were already in the study.

  1. just-solve (control). The agent gets the spec and nothing else.
  2. cohesion-prompt (prompt lever). The same task, plus a plain request for good structure. The delta versus the control isolates what better prompting alone buys you.
  3. impact_gated (tool lever). A neutral prompt, but a change that trips a structural-impact threshold gets one tool-guided refactor and is re-attempted.
  4. metric-in-prompt (kept for reproducibility). This one hands the AI the exact cost formula as its objective.

Condition 4 is the cautionary tale. When you tell the model the precise formula it is scored on, it optimises the proxy instead of the code. It dispersed logic into a pipeline of tiny greenfield classes so the surrounding complexity term collapsed to one and it duplicated code because reuse meant editing a large class the formula punished. The metric went down. The code got worse. Classic Goodhart.

So every lever since then has kept the metric away from the agent. The question that opened up was a different one: the first four signals are all exogenous. Nothing, a prompt, a tool, a formula. What if the quality signal came from the model itself?

The fifth arm: an independent review

The new condition, called reviewed, adds an endogenous signal: the model's own architectural judgment, wired up the way a real pull request review works.

Each checkpoint runs three steps.

  1. The author implements the change with a neutral prompt. Spec only, no formula, no hints about the metric. It optimises the real task.
  2. An independent reviewer critiques the change. This is a separate, fresh session. Its brief is deliberately qualitative and metric-free: is the logic in the right unit, is anything duplicated, is a class or method turning into a catch-all, would a maintainer of this code base find the change natural?
  3. The original author is resumed to fix. If there are findings, the author session is brought back with --resume, so it still has full memory of the change it just made. It is handed the reviewer's notes. It tidies up where it agrees and pushes back where it does not.

The reviewer brief stays qualitative on purpose. Handing the reviewer the cost formula would just re-import the Goodhart gaming from condition 4 through the back door. This arm is a test of whether the model's own taste for structure, prompted only in plain terms, is enough to keep code coherent as it accumulates.

The loop is advisory by construction. It never stops the chain and it makes no intermediate commits. That means the delta versus just-solve measures one thing cleanly: the effect of an independent review loop. And the delta versus impact_gated measures something sharper: the model's own judgment against a metric-driven gate. Endogenous versus exogenous, head to head.

The part I had to get right

The whole experiment rests on a blind-agent rule. Every checkpoint starts stateless. The agent does not carry memory from one change to the next, because if it did, the decay you measure would be tangled up with whatever the agent happened to remember.

Resuming the author looks like it breaks that rule. It does not, and the reason is the load-bearing design decision of this arm. The retained context lives strictly within a single checkpoint. It is torn down at the checkpoint boundary. Checkpoint N+1 still starts blind, exactly like every other arm. And the author and the reviewer never share a session. The only thing that crosses between them is the reviewer's findings as text. So the cross-checkpoint blind condition is preserved, and within a checkpoint the author gets to fix its own work with the context a real engineer would have.

Why this arm matters

Most of the industry conversation about AI code review is about catching bugs. This arm is asking a narrower and, I think, more durable question. Left to accumulate changes, an AI author lets complexity concentrate. Can a second AI, looking only at structure, with no power to run tests and no knowledge of the metric push that back?

Friday, 25 September 2026

Complexity is fixed only the shape moved

Part of a series on how software architecture shapes AI driven code degradation. The previous ten posts published every metric one at a time. This post uses them to make a single argument.

Archived version of record, with the full LaTeX source and figures: doi:10.5281/zenodo.22967550

Tesler was right

Larry Tesler's conservation law says that a problem carries an irreducible amount of complexity. You can move it between the user and the program or between one part of the program and another but you cannot delete it. Fred Brooks says the same thing from the other side. Essential complexity belongs to the problem. Accidental complexity belongs to the solution and only the accidental part is available to be argued about.

This experiment can test that directly because the problem is held fixed. Sixty change requests, identical in both arms, identical in all four conditions. If conservation holds the total amount of complexity in the finished code base should be roughly the same in all eight cells no matter which architecture absorbed it and no matter how the agent was coached.

It is.

amountarchjust-solvecohesiongatedformulaspread, all 8
total_ccSpring65269464664215%
OfficeFloor656681641597
halstead_volumeSpring215,300220,800209,000206,1007%
OfficeFloor216,600221,500214,100207,800
ck_wmc_totalSpring76282676175020%
OfficeFloor738766725679
pmd_cognitive_totalSpring32329329328318%
OfficeFloor319289310268
java_locSpring2,6232,7532,5792,47820%
OfficeFloor2,4512,4922,4182,253

Mean over the last twelve of sixty change requests, averaged over ten runs. Spread is (max minus min) divided by the mean, over all eight cells.

Four independent operationalisations of "amount". Control flow, program vocabulary, weighted methods, and cognitive nesting. They use different theories and different tools. They agree. The finished application carries about six hundred and fifty points of cyclomatic complexity and about two and a half thousand lines of Java and it did not matter which architecture held it and it did not matter what the agent was told.

That is the essential complexity of this problem. It is the floor. No prompt reached under it.

Total cyclomatic complexity across four conditions and two architectures

Eight series that will not separate. This is what a conserved quantity looks like. Click for full size.

Those are totals for the finished system, which is the thing you have to maintain. It is worth also asking what the sixty rules themselves added because the two applications do not start from the same place. Spring's baseline is 413 points of cyclomatic complexity and OfficeFloor's is 361, so a matching total is not automatically a matching amount of work.

complexity added by the sixty rulesjust-solvecohesiongatedformulaspread
Spring total_cc23928523623520%
OfficeFloor total_cc29632128024128%
Spring halstead_volume56,50062,30051,10048,50025%
OfficeFloor halstead_volume58,80063,50056,20050,80022%

Final phase mean minus that run's own checkpoint 1 value, per run, then averaged. Spread is over the four conditions for that architecture.

The conservation claim survives this stricter test and it picks up an honest qualifier on the way through. The spread widens from 7% on totals to between 20% and 28% on what was added, so the amount is steady rather than fixed. And OfficeFloor added more complexity than Spring in every single condition. That is the price of a composed pipeline and it is a real cost visible before any argument about placement begins. OfficeFloor's advantage is not that it carries less. It is that its baseline was lower and it stayed where it was put.

So the only question left is where it goes

If the amount is fixed then every argument about architecture is an argument about arrangement. That is Brooks' point restated as an engineering problem. You cannot negotiate the essential part, so the whole craft lives in the accidental part, which here means one thing: where the six hundred and fifty points of complexity end up sitting.

Arrangement needs its own measures and they are not the same measures as amount. An arrangement metric has to answer "how evenly is this spread" while staying blind to "how much of it there is". Three families do that:

  • Concentration of code. Take every file's share of the total complexity and measure the inequality. ccdist_file_top1 is the share held by the single heaviest file. ccdist_file_hhi is the Herfindahl index, which is the sum of the squared shares. ccdist_file_gini is the Gini coefficient, which is scale free so it sees shape alone and cannot be improved by splitting files.
  • Concentration of change. Hassan's change entropy applied cumulatively from the first commit. cum_change_top1 is the share of all sixty rules' edits that landed in one file. cum_change_entropy_norm is the same information as a spread, where 1.0 is perfectly even. This is the one that speaks directly to maintenance because it describes the file you will be opening again next week.
  • Reach. node_cc_median follows the call graph out from the endpoint and adds up the complexity you have to read to change one rule. ck_cbo_mean and propagation_cost ask the same question through coupling.

Every one of these can move a long way while total_cc does not move at all. That is the property that makes them the right instruments here.

The result: one architecture moved and one did not

Here is the whole finding in one figure. Each row is a metric. Each marker is one condition. The horizontal position is that condition's value divided by the same architecture's own control, so 1.0 means the intervention changed nothing. The number at the end of each bar is the spread across the four conditions.

Amount conserved in both architectures, organisation moved in Spring only

Top panel is amount, where nothing should move. Bottom panel is organisation. Normalising to each architecture's own control is deliberate, because the question here is how far each one moved, not where it started. The absolute levels are in the table below. Click for full size.

The top panel is the conservation result again. Every marker sits on the line for both architectures under every condition.

The bottom panel is the argument. Read the orange rows first. A plain English request for cohesion cut Spring's share of change in one file by two thirds, from 0.344 to 0.110. The formula in the prompt cut its change concentration, cum_change_hhi, by a factor of four. Spring's arrangement is different under every condition, and dramatically so.

Now read the blue rows. They are short. Under the same four interventions, OfficeFloor's arrangement barely registered that anything had been asked of it.

organisation metricSpring spreadOfficeFloor spreadratio
cum_change_top1110%11%10.0
cum_change_entropy_norm24%3%9.2
cum_change_hhi142%20%7.3
mi_mean5%1%5.6
ccdist_file_top187%17%5.3
ck_cbo_mean14%3%4.2
ccdist_file_gini34%8%4.1
total_files37%9%4.0
ccdist_file_hhi106%28%3.7
node_cc_median87%24%3.5
wmcdist_class_hhi100%29%3.4
propagation_cost36%12%2.9
and, for contrast, the amount metrics from the first table
halstead_volume7%6%1.1
total_fns15%16%1.0
total_cc8%13%0.6

Spread is (max minus min) divided by the mean of the four conditions, for that architecture's final phase. Ratio is Spring's spread divided by OfficeFloor's. A ratio near 1 means the intervention moved both equally.

The pattern is clean. On amount the ratio is about 1. Both architectures are equally immovable, which is the conservation law. On organisation, the ratio runs from 3 to 10. One architecture is plastic and the other is not.

Spring's best result is where OfficeFloor started

The interventions worked. That needs saying plainly because it is the strongest thing an opponent of this thesis can say. Spring got structurally better under all three.

The question is what "better" converged on.

metricSpring, controlSpring, bestOfficeFloor, control
cum_change_top10.3440.1100.106
ccdist_file_top10.1970.0780.086
ccdist_file_hhi0.07270.02320.0217
wmcdist_class_hhi0.05540.01910.0182
cum_change_entropy_norm0.7210.9210.886

"Spring, best" is whichever of the three interventions scored best on that metric. OfficeFloor's column is its plain control, with no intervention at all.

Read the right hand column. That is OfficeFloor with no prompt, no tool and no formula. It is the same place Spring arrives after three rounds of coaching.

Spring can be made to distribute its complexity. It has to be asked. OfficeFloor distributes because there is nowhere else for the complexity to go.

Why the two architectures behave differently

A Spring endpoint is a method. A rule is a statement you add to it. Nothing in the framework says where that statement goes, so every rule is a fresh decision made sixty times. Sixty decisions is sixty opportunities for the surrounding instruction to change the answer. That is exactly what the data shows and it is why the same architecture produced four different shapes under four different prompts.

An OfficeFloor endpoint is a graph of wired functions. A rule is a new function and a new edge. The composition is the restriction. There is no version of "add a rule" that concentrates it because the only available move is additive. A prompt asking for cohesion is asking for something the architecture has already done.

Look at what the escape route cost Spring. Under the formula condition Spring's count of container dispatched classes went from 3 to 11.2 per run. Those are advice classes, aspects, filters and entity listeners. They are the framework's own way of running code without anybody calling it. OfficeFloor's count is exactly 1.0 in all four conditions and it never moved.

Container dispatched classes across four conditions

The orange line under the formula condition is complexity leaving the call graph rather than leaving the codebase. Click for full size.

The file count tells the same story from the other end. Asked for structure Spring went from 63 files to 92. OfficeFloor sat between 168 and 184 in every condition because it was already there.

What this does not show

Five things because an argument that only reports its wins is not worth reading.

Function count is not conserved. Of the amount metrics total_fns is the weakest. The rules added between 93 and 168 functions depending on the condition, a spread of 61%. That is what you would expect because "split this into smaller pieces" is an instruction about function count and three of the four conditions were asking for exactly that. Counting functions measures the arrangement as much as the amount. The same is true of pmd_npath_total and halstead_effort, which are published in the amount group but compose super linearly inside a method, so they fall when a method is split. Control flow, volume and weighted methods are the amount metrics that actually hold still.

Function level distribution does not discriminate. ccdist_fn_gini moved 15% in Spring and 14% in OfficeFloor. cogdist_fn_gini moved 16% in both. The inequality of complexity across individual functions is an architecture neutral property and the interventions moved it equally in both arms. The signal lives at the file and class level, which is where the architectural decision actually is. If you only measured function level Gini you would conclude there is no effect here.

Correctness moved in both arms and it moved badly. The strict pass rate, meaning every test of every rule so far green, fell in both architectures under coaching. Spring went from 0.267 under the control to 0.108 under the cohesion prompt and 0.033 under the formula. OfficeFloor went from 0.325 to 0.050 under the cohesion prompt. Both architectures are plastic on correctness. The claim here is about structure only and structural stability did not buy safety. Asking an agent to restructure while it implements is expensive in both worlds.

Read those particular means with more suspicion than the rest of the page because the underlying distribution is bimodal and the mean sits in a gap where almost no run actually lands. In every one of the eight cells most runs finish at 0 or 0.167 and two or three finish near 0.9 and it is those few that hold the average up. Spring's control mean of 0.267 is three runs at 0, five at 0.167 and two at 0.917. The ordering between conditions survives this because the count of runs stuck at 0 moves the same way the mean does, from 3 of 10 under the control to 8 of 10 under the formula. The level does not. No typical run scores 0.267. You can see the whole distribution in the strip beneath the four panels of the strict pass figure, which is what that strip is for.

Two of the three interventions optimise a score this experiment defined. The impact metrics are not evidence for a claim about this experiment because the gated and formula conditions were built to move them. They are reported and they are excluded from the argument above. Every metric in the tables is either a published definition or a whole codebase count.

The handler metrics flatter OfficeFloor and should be read with care. wmc_handler moved 154% in Spring and 135% in OfficeFloor, which looks like a tie. It is not a tie in absolutes. Spring went from 126.5 to 22.6. OfficeFloor went from 1.4 to 5.5. A metric scoped to the entry handler cannot see a pipeline arm so it understates OfficeFloor's real cost. node_cc_median, which follows the calls is the honest version and is the one in the table.

What it means

Tesler's law is usually quoted as a warning about user interfaces. It is a stronger claim than that. Eight independent attempts landed on the same amount of complexity: within 7% of each other on total Halstead volume and within 28% on what the sixty rules added. Against organisation metrics that moved by 100% and more under the same interventions that is a quantity refusing to be argued with. You do not get to remove essential complexity. You get to decide where it lives.

That makes "how is it organised" the only real architectural question and this experiment puts a number on how that question gets answered in each of the two worlds.

In the mutative world the answer comes from outside the code. It came from the prompt, from the review tool, and from the published metric. Change any of those and the architecture changes with it. That is not a bug in Spring. It is what it means to have a framework that permits everything: the structure is whatever the last person to touch it decided and with an AI agent doing the touching sixty times, the structure is whatever you remembered to ask for.

In the additive world the answer comes from the composition itself. Better prompting did not improve OfficeFloor's structure because there was no slack for it to take up. It just kept doing the same thing.

There is a practical version of this. If your structure depends on prompting, reviews and metrics, then it depends on the discipline of whoever is holding them every single time forever. If your structure is a property of the composition it holds when nobody is watching.

Wednesday, 23 September 2026

What it cost and whether it stayed correct

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Cost per change request (USD)

cost_usd  ·  ↓ lower is better  ·  Anthropic API billing, as reported by the agent CLI per turn

Cost per change request (USD)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

What one change request cost to implement, in dollars of model usage. Each change request is a fresh agent session with no memory of the previous one, so this is the whole cost of reading enough of the codebase to understand it, deciding what to do, and writing the code. It is the most direct answer to the question the whole experiment exists to ask, which is whether architecture changes what maintenance costs as a system accumulates rules.

How it is calculated
cost = input tokens · pin + cache reads · pcache + output tokens · pout

Where. The agent CLI reports this figure directly, so the harness does not compute it. The three p terms are the model's published per token prices, which differ: a cache read is billed well below a fresh input token, and an output token well above both. Only the implement turn is counted. The gate's refactor turns and the cold-reader probe are billed to their own separate fields. Each figure plots the mean across the ten independent runs at that change request. The band is one standard deviation across those ten.

In this harness. Read from the agent's terminal JSON for the implement turn at that checkpoint. No modelling and no estimation.

How to read it. A rising line means each successive rule costs more than the last. The codebase is getting more expensive to change. A flat line means the sixtieth rule costs about what the first one did. The gap between the arms at the right hand edge is the compounding penalty for the architecture.

Careful. Cost mixes reading and writing. A run can also get cheaper by giving up. Read it beside the correctness metrics, never alone.

Model inference time per change request

duration_api_ms  ·  ↓ lower is better  ·  agent CLI, API time excluding local tool execution

Model inference time per change request

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How long the model spent thinking and generating for one rule, in milliseconds. It excludes the time spent running Maven, booting a JVM and executing tests. That exclusion matters: wall-clock time is dominated by the build, which is a property of the toolchain and not of the architecture, so wall-clock would bury the signal under noise.

How it is calculated
duration_api_ms = Σrequests in the turn (response received − request sent)

Where. Summed over every model request the implement turn made, which is one per agent step. Local tool execution happens between requests and is therefore not inside any of the intervals. Each figure plots the mean across the ten independent runs at that change request. The band is one standard deviation across those ten.

In this harness. Read from the agent's duration_api_ms for the implement turn. The harness also records duration_ms, the wall-clock figure, which is not plotted for the reason given above.

How to read it. This tracks cost closely, and for the same reason: more context to read and more code to write. It is the independent confirmation that a cost difference is real work rather than a billing artefact.

Cache-read tokens: the comprehension proxy

cache_read_tokens  ·  ↓ lower is better  ·  agent CLI token accounting

Cache-read tokens: the comprehension proxy

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How much existing context the agent had to pull back in to make the change. This is the closest available proxy for the question that actually matters to a team, which is how much of this codebase you have to understand before you can safely touch it. Every checkpoint is a fresh session, so nothing carries over from the last rule. Whatever the agent reads, it reads again from scratch.

How it is calculated
cache_read_tokens = tokens served from the prompt cache during the implement turn

Where. A token is roughly three quarters of an English word, or a few characters of source. The figure counts prompt content the API served from cache rather than re-processing. It includes the harness's own fixed instructions, which are constant across both arms and all four conditions, so the constant part cancels when arms are compared. Each figure plots the mean across the ten independent runs at that change request. The band is one standard deviation across those ten.

In this harness. Read from the agent's token accounting for the implement turn. The harness never uses session resume, so no prior conversation is ever read back. The only thing crossing between checkpoints is the code itself.

How to read it. Rising means the agent has to hold more of the system in its head to add one rule. That is the machine analogue of a developer's ramp-up time, and it is the outcome the concentration metrics are meant to predict.

Careful. It is a proxy. The constant harness prompt is part of the total, so read the trend and the arm gap rather than the absolute level.

Turns taken per change request

num_turns  ·  ↓ lower is better  ·  agent CLI

Turns taken per change request

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many agent steps the implement turn took. One step is one model response plus whatever tools it called. It is a rough measure of how much trial and error the change needed, because a first attempt that compiles and passes ends the turn quickly.

How it is calculated
num_turns = count of agent steps until the implement turn completes

Where. Counted by the agent CLI. A step that only reads a file counts the same as a step that rewrites one. Each figure plots the mean across the ten independent runs at that change request. The band is one standard deviation across those ten.

In this harness. Read from the agent's num_turns for the implement turn.

How to read it. A codebase that fights back produces more turns. Read it with cost, which it partly drives.

Output tokens per change request

output_tokens  ·  · descriptive  ·  agent CLI

Output tokens per change request

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How much text the model generated for one rule. That covers the code it wrote, the edits it issued, and its own reasoning.

How it is calculated
output_tokens = tokens generated by the model during the implement turn

Where. Output tokens are the most expensive of the three billing categories, so this is also the largest single driver of the cost line. Each figure plots the mean across the ten independent runs at that change request. The band is one standard deviation across those ten.

In this harness. Read from the agent's token accounting for the implement turn.

How to read it. Mostly a size signal rather than a quality one. Its real use is as a sanity check that a cheaper arm is not simply doing less work.

Strict pass rate: was everything green after this rule

strict_pass  ·  ↑ higher is better  ·  SWE-CI (arXiv:2603.03823) gate semantics

Strict pass rate: was everything green after this rule

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Whether every selected test was green after the change. That means this rule's own tests plus every earlier rule's tests. The suite is black-box and the agent never sees it. This is the safety headline of the whole experiment, because a rule landed at the cost of breaking two earlier ones is not progress.

How it is calculated
strict_pass(c) = 1 if passed(c) = selected(c), else 0
plotted value = (1/K) Σk=1..K strict_passk(c)

Where. c is the change request, numbered 1 to 60. selected(c) is every test belonging to change requests 1 through c. passed(c) is how many of them were green. K is 10, the number of independent runs, so the plotted value is the fraction of runs that were fully green at that change request.

In this harness. The harness runs the selected Surefire tests after every checkpoint and parses the XML. A checkpoint whose test run crashed is flagged invalid and excluded rather than scored, because a crashed fork reports a successful build with zero selected tests and would otherwise read as the entire prior suite regressing.

How to read it. A line that sags toward the later change requests means the agent is landing new rules while quietly breaking old ones. Compare conditions here before believing any structural improvement. An intervention that improves structure while dropping this line has not made the codebase better.

Careful. Use this, and not the raw regression counts, to ask whether a condition held the suite. A rule broken once and never repaired keeps failing here at every later change request, which is the honest accounting. The transition counts do not show it.

Isolated pass rate: did this rule itself work

iso_pass  ·  ↑ higher is better  ·  harness; the non-regression half of the suite

Isolated pass rate: did this rule itself work

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Whether the change request itself was implemented correctly, ignoring whether it broke anything earlier. Held up against the strict rate, it separates two very different failures: could not do the task, versus did the task and broke something else.

How it is calculated
iso_pass(c) = 1 if every Core, Error and Functionality test of change request c passes, else 0

Where. The acceptance suite is split into four families. Core is the happy path. Error is the rejection and edge-case behaviour. Functionality is hidden behaviour the agent was not told about. Regression is every earlier change request's tests. This metric uses the first three and excludes Regression by construction.

In this harness. Same Surefire parse as the strict rate, filtered to this checkpoint's own three families.

How to read it. A high isolated rate with a low strict rate is the signature of accumulating damage. The agent can still do each new task. It just cannot do it without breaking the last one.

Core pass rate: is the endpoint still alive

core_pass  ·  ↑ higher is better  ·  harness; the happy-path suite

Core pass rate: is the endpoint still alive

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Whether the basic happy path of the endpoint still works. Creating a valid owner and getting a 201 back.

How it is calculated
core_pass(c) = 1 if every Core test of change request c passes, else 0

Where. Core tests are the ones that must never fail. They are a small family per change request, so the denominator is small and the metric is coarse by design.

In this harness. Same Surefire parse, filtered to the Core family.

How to read it. Near 1.0 everywhere is expected. Any visible dip is a serious failure and is worth chasing in that run's own summary rather than here.

Regressions introduced at this checkpoint

regressions  ·  ↓ lower is better  ·  SWE-CI; pass to fail transitions

Regressions introduced at this checkpoint

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many tests were green before this rule landed and red after it. It is a rate of new breakage, not a level of damage. It answers the question of what this one checkpoint broke.

How it is calculated
regressions(c) = | passing(c−1) ∖ passing(c) |

Where. passing(c) is the set of test identifiers green after change request c. The backslash is set difference, so the count is tests in the earlier set but not the later one. A test that was already red before this checkpoint cannot appear here.

In this harness. The harness keeps the pass or fail map per checkpoint and differences consecutive maps by test identifier.

How to read it. Spikes mark the change requests that broke things. It is a per event measure and does not accumulate.

Careful. Two conditions can have near-identical totals here while differing twentyfold in how many tests are standing broken at any moment. This metric does not re-count a rule that was broken earlier and never fixed. For how broken it is right now, use the strict pass rate.

Unintended regressions: the safety signal

true_regressions  ·  ↓ lower is better  ·  SWE-CI, with the harness's intended and unintended split

Unintended regressions: the safety signal

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Regressions the change request did not ask for. Some change requests deliberately revise an earlier rule, and when they do, the earlier rule's tests are supposed to change. Those do not count here. What is left is the cleanest breakage signal on the page, because it means the agent broke something it was never asked to touch.

How it is calculated
true_regressions(c) = | (passing(c−1) ∖ passing(c)) ∖ intended(c) |

Where. intended(c) is the set of earlier tests that change request c declared it would change, which the checkpoint plan records as its mutates list. For a purely additive change request intended(c) is empty, so every regression is a true one.

In this harness. The mutates declaration lives in checkpoints.yaml beside the change request, so the intended set is fixed before any run starts and cannot be fitted to a result afterwards.

How to read it. This is the number that means it broke something it was not asked to break. Flat at zero is the target.

Normalized change

normalized_change  ·  ↑ higher is better  ·  SWE-CI (arXiv:2603.03823)

Normalized change

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

A single score for how much the change moved the suite toward its target. It is positive for progress and negative for regression, and deliberately asymmetric: breaking things is scored against a different denominator from fixing them, so a small amount of breakage in a large suite is not lost in rounding.

How it is calculated
NC(c) = (passed − base) / (target − base)  if passed ≥ base
NC(c) = (passed − base) / base  if passed < base

Where. base is how many tests were passing before this change request. passed is how many are passing after it. target is the total number selected, which is the best achievable. The result lies in the range −1 to 1. The improvement branch is progress toward what was still missing. The regression branch is loss measured against what already worked, which is why it bites harder.

In this harness. Computed in harness/correctness.py from the same pass or fail maps as the other correctness fields. Degenerate denominators are guarded: a zero baseline scores −1 on the regression branch and a zero gap scores 1 on the improvement branch.

How to read it. A compact per checkpoint verdict. Its use is spotting which phase of the sixty rules a condition started losing ground in.

Cold-reader recall: can a fresh agent still find the rules

probe_recall  ·  ↑ higher is better  ·  harness read-only comprehension probe

Cold-reader recall: can a fresh agent still find the rules

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

A fresh agent with no history is dropped into the codebase and asked what business rules the create-owner endpoint enforces. This is the fraction it finds. It is the one metric here that measures comprehension directly rather than by proxy, which is why it is worth the cost of running it.

How it is calculated
probe_recall = | rules found | / | rules implemented so far |

Where. rules implemented so far is the change requests landed up to that checkpoint, which is known exactly because the plan is fixed. rules found is how many of them the probe named. The probe runs in its own session with no tools that can modify anything.

In this harness. Run periodically rather than at every checkpoint, which is why the line is sparse. The probe's own cost and tokens are recorded separately from the implement turn so they never contaminate the cost line.

How to read it. This is the comprehension outcome the whole experiment is about. Not whether the code is complex but whether a newcomer can find out what it does.

Careful. Sparse by design. Only a fraction of checkpoints carry a probe, so the line has few points and each one is noisier than a dense metric.

How much complexity there is

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Total cyclomatic complexity

total_cc  ·  = prediction: no arm difference  ·  McCabe 1976, via lizard

Total cyclomatic complexity

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The total number of independent paths through the whole application. Every branch point in every function, summed. It is the standard answer to how much decision-making a codebase contains, and it is the primary test of Tesler's conservation law here: the sixty change requests demand a certain number of decisions, and no architecture can delete them.

How it is calculated
total_cc = Σf ∈ F CC(f)

Where. F is every function in production Java at that checkpoint's commit. CC(f) is the cyclomatic complexity of function f, as reported by lizard: one plus the number of decision points, counting if, for, while, case, catch, the ternary operator, and each && or ||. The population is production Java only, meaning the files matched by the arm's source_globs at that checkpoint's commit. Tests are excluded. OfficeFloor's YAML wiring is counted in a separate pool and is never folded into a Java denominator, because mixing them would let a wiring-based architecture dilute any per-line metric.

In this harness. lizard parses every matched file and reports one record per function. The harness sums the cyclomatic complexity column. A parser that fails on a file yields no functions for it, which would make that file look free, so parser_selftest must pass before a run or an analysis is trusted.

How to read it. The prediction is that the two lines sit on top of each other. If a condition lowers this, it either skipped work or pushed logic somewhere the parser cannot read. Check the validity group before celebrating.

Total cognitive complexity

pmd_cognitive_total  ·  = prediction: no arm difference  ·  Campbell / SonarSource 2018, via PMD

Total cognitive complexity

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

A complexity measure built around how hard code is for a human to follow, rather than how many paths it has. Nesting is penalised heavily. A long flat sequence of independent checks is forgiven. That makes it the one measure in this group that could separate a deeply nested god method from a long flat dispatch, which is exactly the distinction cyclomatic complexity is blind to.

How it is calculated
cognitive(f) = Σs ∈ structures(f) (1 + nesting(s))
pmd_cognitive_total = Σf ∈ F cognitive(f)

Where. structures(f) are the flow-breaking constructs in f: if, else if, else, switch, each loop, each catch, each ternary, and each run of mixed && or || operators. nesting(s) is how many enclosing structures s sits inside. Some constructs take the increment without contributing nesting, which is what forgives a flat sequence: ten sibling if statements score 10, while ten nested ones score 55.

In this harness. PMD computes it with its own Java parser. The harness spawns PMD once per checkpoint for both of its rulesets and splits the merged report, because a JVM launch per ruleset would dominate the analysis time.

How to read it. Agreement with total cyclomatic complexity means the conservation result is not an artefact of how branches are counted. Disagreement would mean one arm's complexity is more deeply nested, which is a real finding about shape rather than amount.

Total NPath complexity

pmd_npath_total  ·  = prediction: no arm difference  ·  Nejmeh 1988, via PMD

Total NPath complexity

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The number of acyclic execution paths through a function. It is cyclomatic complexity's multiplicative cousin. Where McCabe adds across sequential branches, NPath multiplies, which is closer to the number of distinct behaviours a test suite would have to cover. Two sequential if-statements are 3 by McCabe and 4 by NPath. Ten of them are 11 versus 1024.

How it is calculated
NP(if) = NP(cond) + NP(then) + NP(else)
NP(seq) = NP(s1) × NP(s2) × …
pmd_npath_total = Σf ∈ F NP(body of f)

Where. The recursion is over the statement tree. Statements in sequence multiply, which is where the explosion comes from. A branch adds its arms. A statement with no control flow has NP 1. Loops and switch have their own rules in Nejmeh's original paper, which PMD implements.

In this harness. PMD, from the same single spawn as cognitive complexity.

How to read it. Because it multiplies, NPath detonates when decisions pile up in the same method. A conservation result that survives NPath is a strong one.

Careful. The scale is enormous and is dominated by whichever single method has the most sequential branches. Read the shape of the line, not its value.

Total Halstead volume

halstead_volume  ·  = prediction: no arm difference  ·  Halstead 1977

Total Halstead volume

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Program size measured by vocabulary rather than by control flow. It asks how many distinct operators and operands there are and how often they appear. It has no concept of a branch at all, which is what makes it valuable here: if it agrees with the cyclomatic total, the conservation result is not a property of one family of measure.

How it is calculated
V = N · log2 η
where η = η1 + η2 and N = N1 + N2

Where. η1 is the number of distinct operators and η2 the number of distinct operands, so η is the vocabulary. N1 and N2 are the total occurrences of each, so N is the program length. Volume is therefore length times the bits needed to name one vocabulary item.

In this harness. The harness tokenises each file itself in placement.halstead. Symbolic operators and the control-flow keywords count as operators. Identifiers, numbers and type keywords count as operands. Comments are stripped. Every string and character literal is replaced by one placeholder operand before tokenising, so an arm cannot move its Halstead score by writing longer error messages. File volumes are summed.

How to read it. The fourth independent way of asking how much is there. Treat the units as arbitrary and compare the two lines.

Total Halstead effort

halstead_effort  ·  = prediction: no arm difference  ·  Halstead 1977

Total Halstead effort

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Halstead's own estimate of the mental work needed to write or understand the program. It is volume multiplied by difficulty, where difficulty grows as a small vocabulary of operands gets reused many times. It is more sensitive than volume, and it moves for reasons volume does not.

How it is calculated
D = (η1 / 2) · (N2 / η2)
E = D · V

Where. D is difficulty. The first factor is half the operator vocabulary. The second is the average number of times each distinct operand is used, so a function that keeps reusing the same few variables scores as harder. V is the volume above. Effort is their product, summed over files.

In this harness. Computed in the same tokenising pass as volume.

How to read it. Read it as a more sensitive volume. A gap here with no gap in volume means one arm reuses its operands more heavily.

Total weighted methods per class (independent parser)

ck_wmc_total  ·  = prediction: no arm difference  ·  Chidamber & Kemerer 1994, via the CK tool (Aniche 2015)

Total weighted methods per class (independent parser)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same total-complexity question, computed class by class, by a completely different tool with a completely different parser. This is a cross-check rather than a finding. Every other complexity number on this page comes from lizard. If a parser quirk were driving the conservation result, this is where it would show up.

How it is calculated
WMC(C) = Σm ∈ methods(C) CC(m)
ck_wmc_total = ΣC ∈ classes WMC(C)

Where. C ranges over every class the CK tool can parse. Chidamber and Kemerer left the per-method weight open. CK uses cyclomatic complexity, which is the conventional choice and matches the lizard side, so the two totals are directly comparable.

In this harness. CK is a Java program that parses source with Eclipse JDT, which is a full compiler front end. lizard uses its own lightweight parser. The harness runs CK once per checkpoint over the arm's source directories and reads its per-class CSV.

How to read it. Agreement with total cyclomatic complexity is the result you want, and it is a boring one. Disagreement means one of the two parsers is failing to read some file, which would make that file look perfect.

Maintainability Index, average file

mi_mean  ·  ↑ higher is better  ·  Coleman et al. 1994, in the SEI and Visual Studio rescaling

Maintainability Index, average file

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

A composite index built from Halstead volume, cyclomatic complexity and lines of code, rescaled so that higher is more maintainable. It is the number a reviewer expects to see, and it combines all three families of measure, so it is a useful single check on whether aggregate maintainability differs at all before any placement argument is made.

How it is calculated
MI(file) = max(0, (171 − 5.2 ln V − 0.23 CC − 16.2 ln LOC) · 100/171)
mi_mean = mean over files of MI(file)

Where. V is that file's Halstead volume, CC its summed cyclomatic complexity and LOC its line count. The constants are Coleman's, fitted in 1994. The · 100/171 factor and the floor at zero are the later SEI rescaling into a 0 to 100 range, which is the form most tools report. It is computed per file and then averaged, not over the codebase as one blob.

In this harness. Computed in placement.maintainability_index from the harness's own Halstead pass and lizard's per-function complexity, grouped by file.

How to read it. It is location-blind by construction, which is precisely why it must not be the thesis statistic. It cannot see concentration.

Careful. Averaging over files rewards fragmentation. An architecture that adds many small healthy files raises its own average without improving anything. Read it with the worst-file version below, which cannot be diluted.

Maintainability Index, worst file

mi_min  ·  ↑ higher is better  ·  Coleman et al. 1994

Maintainability Index, worst file

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The Maintainability Index of the single worst file in the codebase. Unlike the average it cannot be diluted by adding good files, which makes it the version worth quoting.

How it is calculated
mi_min = min over files of MI(file)

Where. Same per-file MI as above. The minimum is taken over every production Java file with a positive volume and line count.

In this harness. Same pass as the mean.

How to read it. A falling line means the worst thing in the codebase is getting worse, whatever else is happening. That is usually the thing a maintainer will actually meet.

Function count

total_fns  ·  · descriptive  ·  lizard

Function count

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many functions exist in production Java. This is descriptive, and it is also the scale correction for several other metrics. A distributed architecture should climb here, because that is the mechanism it works by, not a finding about it.

How it is calculated
total_fns = | F |

Where. F is the same function population as the cyclomatic total. The population is production Java only, meaning the files matched by the arm's source_globs at that checkpoint's commit. Tests are excluded. OfficeFloor's YAML wiring is counted in a separate pool and is never folded into a Java denominator, because mixing them would let a wiring-based architecture dilute any per-line metric.

In this harness. Count of lizard's function records at the checkpoint commit.

How to read it. It matters because several concentration metrics are not scale-free. An arm with more units scores better on those for free. This line is how you check whether a concentration gap is a difference in shape or just a difference in count.

File count

total_files  ·  · descriptive  ·  lizard

File count

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many production Java files exist. The same role as the function count, one level up, and the denominator for every file-based concentration index on this page.

How it is calculated
total_files = | { file(f) : f ∈ F } |

Where. file(f) is the path the function was found in. Only files containing at least one parsed function are counted, which is worth knowing: a file the parser cannot read does not appear here either.

In this harness. Distinct file paths among lizard's function records.

How to read it. Read it beside any HHI or top-share figure. Those are not scale-free, so a rising file count lowers them for free.

Package count

total_packages  ·  · descriptive  ·  harness

Package count

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many Java packages the production code spans. It says whether new rules got their own home in the source tree or were filed into existing ones.

How it is calculated
total_packages = | { package(f) : f ∈ F } |

Where. package(f) is derived from the file's path below the source root, with the file name removed, which is the package a Java file in a conventional layout declares.

In this harness. Derived from file paths in placement._package_of rather than by reading package statements, so it reflects the directory structure a developer navigates.

How to read it. Descriptive. It is context for the package-level concentration index.

Production Java lines of code

java_loc  ·  · descriptive  ·  lizard nloc, summed

Production Java lines of code

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Total production Java, excluding tests and excluding the YAML wiring. The raw size line. Every ratio metric on this page has this or a close relative as its denominator, so a surprising ratio is often a size story.

How it is calculated
java_loc = Σf ∈ F nloc(f)

Where. nloc(f) is lizard's count of non-comment, non-blank lines in the function. Note that this sums function bodies, so file-level declarations, imports and field initialisers are not included. The population is production Java only, meaning the files matched by the arm's source_globs at that checkpoint's commit. Tests are excluded. OfficeFloor's YAML wiring is counted in a separate pool and is never folded into a Java denominator, because mixing them would let a wiring-based architecture dilute any per-line metric.

In this harness. Summed from lizard's per-function records. The YAML pool is counted into yaml_loc separately and the two are never added together.

How to read it. Both arms should grow. The interesting question is whether one grows faster for the same sixty rules, which would mean it needs more code to express the same behaviour.