Wednesday, 23 September 2026

The price of distributing

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Call hops from a handling step

indirection_median  ·  ↓ lower is better  ·  harness breadth-first search over the call graph

Call hops from a handling step

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many calls deep you have to go, from a handling step, to reach the code that does the work. Every hop is a file a developer has to open. This is the sceptic's objection made numeric: you did not remove the complexity, you buried it behind indirection.

How it is calculated
depth(k) = length of the shortest call path from any node to k
indirection_median = mediank ∈ reached depth(k)

Where. reached is every method reachable from the handling nodes. The handling nodes themselves have depth 0. Depth is assigned by breadth-first search, so it is the shortest path, not the longest. The harness also records how many methods were reached, as indirection_reached, which is the denominator.

In this harness. placement.indirection_stats, sharing the same call index as the comprehension metrics.

How to read it. Expected to favour the concentrated arm. Everything in one method is zero hops away. This is Brooks' accidental complexity as a number, and it is the honest price of decomposition.

Careful. It understates a pipeline architecture. Depth is measured from the handling nodes, so a pipeline's twenty wired steps are all at depth 0 while a single-handler arm has exactly one depth-0 method. Worse, a pipeline's hops between steps are declared in YAML and dispatched by the container, so they are not Java calls and do not appear here at all. The honest reading of a pipeline arm is this depth plus its step count. Quoting this column alone would let the arm with the most indirection report the least.

Deepest call chain

indirection_max  ·  ↓ lower is better  ·  harness breadth-first search over the call graph

Deepest call chain

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The longest shortest-path from a handling step to any reachable method. It is the worst-case navigation cost for one change.

How it is calculated
indirection_max = maxk ∈ reached depth(k)

Where. Same depth assignment as the median. Because every depth is a shortest path, this is the eccentricity of the reachable set rather than the length of the longest walk.

In this harness. Same single pass.

How to read it. A rising maximum with a flat median means one long tail appeared rather than a general deepening.

Share of reached methods more than two hops away

indirection_deep_share  ·  ↓ lower is better  ·  harness breadth-first search over the call graph

Share of reached methods more than two hops away

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Of everything the request path reaches, what fraction is far enough away that you would not find it by reading the handler. It is a distribution-shape measure rather than an extreme: not how deep the deepest is, but how much of the system lives out in the far field.

How it is calculated
indirection_deep_share = | { k : depth(k) > 2 } | / | reached |

Where. The threshold of 2 is a choice, and it is the only arbitrary constant in this group. It is meant as “further than the handler and the thing it obviously calls”.

In this harness. Same single pass.

How to read it. Expected to be a counter-signal, like the rest of this group.

Coupling between objects, average class

ck_cbo_mean  ·  ↓ lower is better  ·  Chidamber & Kemerer 1994, via CK

Coupling between objects, average class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many other classes a class depends on. Splitting one class into six creates coupling between the six that did not exist before, so this is expected to be a counter-signal.

How it is calculated
CBO(C) = | { D ≠ C : C references D or D references C } |
ck_cbo_mean = mean over classes of CBO(C)

Where. A reference is any use of the other class: a field type, a parameter type, a local variable, a method call, a thrown exception. CK counts distinct classes, not distinct references, so calling one class fifty times still counts 1. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. If a distributed architecture keeps this flat while adding units, that is a real result in its favour, because it is the objection you would most expect to land.

Coupling between objects, worst class

ck_cbo_max  ·  ↓ lower is better  ·  C&K 1994, via CK

Coupling between objects, worst class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The single most entangled class in the codebase.

How it is calculated
ck_cbo_max = max over classes of CBO(C)

Where. Same CBO definition as the mean.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Where the mean is diluted by many small classes, the maximum is not. A rising maximum with a flat mean means one class is becoming the hub.

Fan-out, average class

ck_fanout_mean  ·  ↓ lower is better  ·  CK

Fan-out, average class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many other classes a class calls out to. It is the directional half of coupling: not how entangled a class is, but how much it depends on others.

How it is calculated
fanout(C) = | { D : C references D } |
ck_fanout_mean = mean over classes of fanout(C)

Where. Unlike CBO this counts only outgoing references. CK records the incoming direction separately as fan-in, which the harness also stores.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Rising fan-out with flat fan-in is the orchestrator shape: a class that coordinates rather than one that is depended upon.

Response for a class, average

ck_rfc_mean  ·  ↓ lower is better  ·  Chidamber & Kemerer 1994, via CK

Response for a class, average

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many distinct methods could end up executing in response to one message to this class. That is its own methods plus everything they call. It is roughly the size of the behaviour you have to consider when you call into a class, which makes it both a testability and a comprehension measure.

How it is calculated
RFC(C) = | M(C) ∪ ⋃m ∈ M(C) R(m) |
ck_rfc_mean = mean over classes of RFC(C)

Where. M(C) is the methods declared by C. R(m) is the set of methods invoked by m. CK uses the one-level version, which is the common implementation: it counts methods called directly by the class's own methods and does not recurse.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. It rises with both concentration and indirection, which makes it a useful tiebreaker between those two stories.

Response for a class, worst

ck_rfc_max  ·  ↓ lower is better  ·  C&K 1994, via CK

Response for a class, worst

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The class with the largest response set. The worst thing in the codebase to call into.

How it is calculated
ck_rfc_max = max over classes of RFC(C)

Where. Same one-level RFC definition as the mean.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Read it beside the handler's own RFC below. If the worst class in the codebase is the handler, that is the thesis. If it is something else, say so.

Response set of the handler class

ck_handler_rfc  ·  ↓ lower is better  ·  CK, pinned to the handler class

Response set of the handler class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How much behaviour is reachable from one call into the endpoint's class. It is closely related to the whole-path complexity in the comprehension group, but computed by a different tool, in units of methods rather than branches, and only one level deep.

How it is calculated
ck_handler_rfc = | M(H) ∪ ⋃m ∈ M(H) R(m) |

Where. H is the handler class. One level only, so unlike the comprehension walk this does not follow the call graph transitively.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. A second opinion on the handler's comprehension load, from a tool with a different parser and a different definition.

Coupling of the handler class

ck_handler_cbo  ·  ↓ lower is better  ·  CK, pinned to the handler class

Coupling of the handler class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many other classes the endpoint's own class depends on. A handler that collects dependencies is a handler that is accumulating responsibilities.

How it is calculated
ck_handler_cbo = | { D ≠ H : H references D or D references H } |

Where. Same CBO definition, scoped to the handler class.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. In a pipeline architecture the dependencies move to the wiring, which is exactly the kind of relocation the framework-dispatch caveat is about. A very low figure here for an arm implementing sixty rules is a prompt to check the validity group.

Law of Demeter violations

pmd_demeter_violations  ·  ↓ lower is better  ·  Lieberherr & Holland 1989, via PMD

Law of Demeter violations

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Places where code reaches through one object to get at another, as in a.getB().getC().doThing(). It is the classic symptom of a class knowing too much about its neighbours' internals, and it is another externally defined detector with thresholds nobody here chose.

How it is calculated
pmd_demeter_violations = count of method calls whose receiver is not this, a parameter, a locally created object, or a field of this

Where. The law says a method may only call methods on: itself, its own parameters, objects it created, and its own fields. PMD's rule reports one violation per offending call site, so a single long chain can contribute several.

In this harness. PMD's LawOfDemeter rule, from the same single spawn as the complexity measures.

How to read it. It tends to rise with distribution, so treat it as a counter-signal. It is also a famously noisy rule, so read the trend rather than the absolute count.

Propagation cost

propagation_cost  ·  ↓ lower is better  ·  MacCormack, Rusnak & Baldwin 2006

Propagation cost

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

If you change a random file, what fraction of the codebase could feel it? It is the density of the transitive closure of the file dependency matrix, and it is the one whole-architecture coupling number here with real pedigree in the modularity literature.

How it is calculated
propagation_cost = Σi=1..n | reach(i) | / n2

Where. n is the number of production Java files. reach(i) is the set of other files reachable from file i by following call edges transitively, so a file does not count itself. The numerator is therefore the number of ordered reachable pairs, and dividing by n2 gives the expected fraction of the system a random change can touch.

In this harness. placement.propagation_cost collapses the method call graph to a file graph, dropping self-edges, then runs a depth-first reach from every file.

How to read it. Lower means better modularised. Use the between-arm comparison at the same change request and nothing else.

Careful. Two load-bearing caveats. First, the n2 denominator rewards having more files, so an arm that splits the same code over more files scores lower for free. Second, it inherits a conservative call resolver that drops every edge it cannot prove. MacCormack reports 10 to 60 percent for real systems, so a value an order of magnitude below that means edges are missing rather than that the design is exceptional. Never quote the absolute value.

Files reachable from a typical file

propagation_fanout_median  ·  ↓ lower is better  ·  MacCormack et al. 2006, unnormalised

Files reachable from a typical file

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The raw count behind propagation cost. From one file, how many files can you reach by following dependencies. Because it is not normalised it cannot be improved by adding files, which makes it the honest version.

How it is calculated
propagation_fanout_median = mediani=1..n | reach(i) |

Where. Same file graph and same transitive reach as the propagation cost, with no n2 division. The harness also records the maximum and the file count.

In this harness. Same single pass as the propagation cost.

How to read it. If the normalised and unnormalised numbers tell different stories, the difference is packaging rather than structure. That comparison is the reason both are reported.

Duplication and boilerplate

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Verbosity: duplicated and anti-pattern lines

verbosity  ·  ↓ lower is better  ·  SlopCodeBench (arXiv:2603.24755) Eq. 4

Verbosity: duplicated and anti-pattern lines

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The share of the codebase that is either copy-pasted from elsewhere in the same codebase or matches a known anti-pattern. This is the metric that catches the cheap way to score well on everything else on this page, which is to copy the logic into a new small class instead of factoring it out.

How it is calculated
verbosity = | clone_lines ∪ pattern_lines | / LOC

Where. clone_lines and pattern_lines are sets of (file, line number) pairs, not counts, which is what makes the union meaningful. The union rather than the sum matters: a line that is both duplicated and an anti-pattern is charged once. LOC is the production Java line count. The result can exceed 1 in principle, because the LOC denominator counts function bodies while the line sets are gathered over whole files.

In this harness. metrics.verbosity. Clones come from jscpd, anti-patterns from PMD or ast-grep depending on the run's own config snapshot, so an old run replays with the detector it actually used. If one detector cannot run, the metric is computed from the other and the harness records which halves ran.

How to read it. One of the four conditions produced exactly the failure this metric exists to catch. Read it beside the placement group, not on its own.

Duplicated lines

verbosity_clone_lines  ·  ↓ lower is better  ·  jscpd clone detection

Duplicated lines

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The count of lines that appear as a near-identical block somewhere else in the codebase. It is the copy-paste half of verbosity, reported on its own so the two halves can be told apart.

How it is calculated
verbosity_clone_lines = | clone_lines |

Where. clone_lines is the set of (file, line) pairs jscpd reports as belonging to a duplicated block. Both copies of a clone are counted, because both are lines a maintainer has to keep in step. jscpd's minimum block size is what decides whether a short repeated idiom counts, and it is pinned in tools/package.json so the threshold cannot drift between runs.

In this harness. jscpd at a pinned version, over the arm's source directories. quality_selftest fails closed if the pinned version is not what is installed.

How to read it. A rising line here while the structural metrics improve is the signature of fragmentation masquerading as decomposition.

Anti-pattern lines

verbosity_pattern_lines  ·  ↓ lower is better  ·  PMD ruleset, or the ast-grep rules in astgrep-rules/

Anti-pattern lines

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The count of lines matching structural anti-patterns. The rules are defined over the syntax tree rather than over text, so they survive reformatting and renaming, which makes them harder to game than a text-based lint.

How it is calculated
verbosity_pattern_lines = | pattern_lines |

Where. pattern_lines is the set of (file, line) pairs flagged by the configured detector. With PMD the ruleset is pmd-rules/java-wasteful.xml. With ast-grep it is the rules in astgrep-rules/. Which one applies is read from the run's own config snapshot.

In this harness. From the same single PMD spawn as the complexity measures, filtered to the wasteful ruleset's rule names. The quality gate routes through the same choice, so gate and metric can never disagree about what a smell is.

How to read it. The other half of verbosity. If this moves and the clone count does not, the agent is writing fresh bad code rather than copying old code.

Validity Guards

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Framework-dispatched classes: the call-graph escape counter

container_total  ·  · descriptive  ·  harness; counts advice, aspects, filters, entity listeners and validators

Framework-dispatched classes: the call-graph escape counter

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many classes are invoked by the framework rather than by an ordinary method call. This is not a finding. It is a lie detector for every other metric on this page that follows a call graph.

How it is calculated
container_total = advice + aspect + filter + entity_listener + validator

Where. Each term counts classes carrying the corresponding framework hook: advice is @ControllerAdvice or @RestControllerAdvice; aspect is @Aspect; filter is a servlet Filter or a HandlerInterceptor; entity_listener is @EntityListeners or a @PrePersist callback; validator is a ConstraintValidator. The harness stores each term separately as well as the total.

In this harness. placement.container_dispatch scans the production Java files for those annotations and interfaces at the checkpoint commit.

How to read it. Not a finding. A lie detector. No call-graph walk can see a class the container dispatches, so every comprehension metric on this page is measuring a shrinking fraction of the code wherever this line rises. In one condition, chains ended with eighteen advice classes and a handler containing the stock upstream body. The whole-path complexity read 3 for a codebase implementing all sixty rules. If this line has moved off its baseline for a condition, discount that condition's comprehension numbers rather than the thesis.

Invalid test gates

gate_invalid  ·  ↓ lower is better  ·  harness resilience layer

Invalid test gates

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Checkpoints where the test run itself failed in a way that would otherwise be scored as a mass regression. A crashed Surefire fork is the usual cause. It reports a successful build and zero selected tests, which naively scores as the entire prior suite regressing at once.

How it is calculated
gate_invalid(c) = 1 if build_ok ∧ selected(c) = 0 ∧ neighbours select many, else 0

Where. The signature is the conjunction: the build succeeded, no tests were selected, and the surrounding checkpoints select dozens. A checkpoint flagged this way is excluded from correctness scoring and retried rather than recorded as a failure.

In this harness. Detected in harness/correctness.py, which blanks every correctness field and sets this flag. analyze also repairs old captures by the same signature, which is why any correctness number produced before the fix must be recomputed rather than trusted.

How to read it. Should be flat at zero. It exists because an unflagged crashed fork once read as 143 regressions where the true figure was 0. The biggest risk to a study like this is infrastructure, not statistics.

Tests selected per checkpoint

total_selected  ·  ↑ higher is better  ·  harness test selection

Tests selected per checkpoint

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many black-box acceptance tests ran at this checkpoint. That is this rule's own tests plus every prior rule's. It rises by construction as rules accumulate, and that rise is what makes the later change requests harder than the early ones.

How it is calculated
total_selected(c) = Σj=1..c | tests(j) |

Where. tests(j) is the acceptance tests belonging to change request j, across all four families. Selection is by change request number, so it is deterministic and identical in both arms.

In this harness. The harness selects by checkpoint from acceptance/ and runs them with Surefire after every checkpoint.

How to read it. A smooth rise is correct. A sudden drop to zero at a checkpoint whose neighbours pass dozens is the crashed-fork signature above.

Classes seen by the independent parser

ck_classes  ·  · descriptive  ·  CK tool

Classes seen by the independent parser

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many classes the second parser could read. A parser that fails on a file produces no metrics for it, which makes that file look perfect, so this is the cheapest available check that both tools are seeing the same codebase.

How it is calculated
ck_classes = number of rows in CK's per-class output

Where. CK counts real class declarations, so nested and inner classes appear as their own rows. That is why this figure normally sits above the file count rather than equal to it.

In this harness. Row count of CK's class CSV at the checkpoint commit.

How to read it. Compare the shape of this line against the file count. A divergence means CK started failing on something, and every CK metric on this page is then measuring less than it claims.

Tuesday, 22 September 2026

We asked nicely and it cost 23 percent

A series on how software architecture shapes AI driven code degradation. This is the plain version of a result. The paper has the numbers and the caveats.

Quick recap of the experiment. An AI agent gets sixty change requests, one after another, all landing on the same REST endpoint. Add a validation rule. Change how phone numbers are stored. Add a duplicate check. The full test suite runs after every single one.

We run that on a normal Spring codebase. We run the same sixty changes on the same application built with OfficeFloor. Then we look at the wreckage.

The Spring result has always been the same. The create owner handler grows. Rule after rule lands in the same method. By the end it is the biggest thing in the codebase.

Last time we put our scoring formula straight into the prompt. We told the AI exactly how it was being measured. It optimised the score beautifully. It also scattered the rules into thirty one new classes, wrote more duplicated code, and in two runs hid every rule somewhere our tooling could not see. Worse, it stopped keeping earlier rules working. The whole suite was green after only 44% of changes, down from 79%.

That left one obvious question hanging. Was that because we showed it the metric? Or would any instruction about structure have done the same damage?

So we ran it again. This time we just asked nicely.

The whole intervention is one paragraph

No formula. No mention of any metric. No mention of the experiment. Just the paragraph a senior engineer might add to a ticket.

As you do, keep the code well structured: put each piece of
logic where it belongs, in a small unit with a single
responsibility, and reuse existing code instead of copying
it. Do not let any one class or method grow into a catch-all
that accumulates unrelated logic.

Everything else stayed identical. Same sixty changes. Same tests. Same model. Ten independent runs per codebase. Nothing checked the AI's work and made it try again. The paragraph is the entire change.

It worked

The controller file starts at 203 lines. Here is how many lines got added to it over sixty rules.

Spring controller, lines added over 60 rulesTen runs
Normal prompt604 to 962
Asked nicely36 to 273
Told the formula2 to 44

One paragraph of plain English cut the god method to about a fifth of its size. Ask the codebase which class is heaviest at the end and the answer changes too. Under the normal prompt it is the controller in nine runs out of ten. After the paragraph it is the Owner entity in ten runs out of ten. That class is heavy because it holds getters and setters. It does not hold decisions.

How hard is it to read the create path at the end? We add up the complexity of everything reachable from the endpoint. That number goes from 201 down to 129. Every single run landed below the worst normal run.

So the short version is that asking works. You do not need to show the AI your metric to get most of the benefit of it.

It was not free

This is the part that surprised us, and it is the reason this post has the title it has.

Per run of 60 changes, SpringNormal promptAsked nicely
What the agent cost$78$97
Tool turns it took1,5221,844
Wall clock4.0 hours4.6 hours

The bill went up 23%. Roughly five extra tool calls per change. The OfficeFloor side went up 17%.

Now compare that to the formula version. That one cost nothing extra at all. Same money, same turns.

That contrast is worth sitting with. Optimising arithmetic is cheap. Exercising judgement is not. When you ask an agent to think about where code belongs, it reads more, looks around more and edits more. You pay for all of it.

One thing we checked, because it is the obvious follow up. Within the ten runs, did the ones that spent more end up better structured? No. Spending more did not buy more. The 23% is the price of the instruction. It is not a dial you can turn.

The code still moved rather than shrank

We have a check that ignores all of our metrics. It takes the finished codebase, finds every line the run changed, works out which method that line now sits in, and adds up how complicated all those methods are. It just asks how much logic exists and where it lives.

Spring, after 60 rulesNormal promptAsked nicelyTold the formula
Total logic written282326268
...sitting in brand new files46205218
...sitting in files that already existed23612150
Number of files involved195841
Duplicated lines at the end2,2302,8202,510

Look at the first row. Sixty business rules are sixty business rules. The logic did not get smaller. It got slightly bigger, because small classes need declarations, constructors and call sites.

The next two rows are the same story as last time. The logic moved out of the files that existed and into files the AI created. Three times as many files as the normal run.

Then look at the last row. Duplication went up by about a quarter. Remember the paragraph we wrote. It contains the words "reuse existing code instead of copying it". That is the one explicit request in it. It is the one thing the run did not deliver.

We do not have a proven reason for that. Our best guess is that the other instruction won. Keep units small, and you end up with a lot of small units. Finding the right existing helper among sixty of them is harder than writing a fresh one. If you have ever worked in a codebase with four slightly different StringUtils classes, you have seen a human do the same thing.

The good news, and it is genuinely good

This run created about 41 new classes per run. The formula run created about 31. So more classes. The interesting bit is what kind.

New classes per run, SpringAsked nicelyTold the formula
Proper injected Spring beans167
Static utility holders1621

This matters more than it looks. A static utility class is a class with one static method and nothing else. On our score it is perfect, because there is no surrounding class weight to pay for. In a real Spring codebase it costs you things you run into fast. You cannot inject anything into it. You cannot swap it in a test. Spring cannot wrap it in a transaction or a proxy.

Told the formula, the AI mostly wrote static holders. That is the cheapest possible answer to "make this number small". Asked in English, it wrote real beans about half the time. That is an actual answer to "put this where it belongs".

And nothing vanished this time. Last run, two of the ten Spring runs moved every rule into interceptors that Spring calls for you. Our comprehension metric follows method calls, and nothing calls an interceptor, so those runs simply disappeared from the measurement. That did not happen here. The count of framework invoked classes stayed near the application's baseline of three, with a worst run of eight. Under the formula it averaged thirteen, with a worst run of thirty three. When the numbers got better this time, the code actually got better.

The part that actually matters

Here is the correctness result across all three prompts.

SpringNormalAsked nicelyTold the formula
Implemented the rule it was asked forevery timeevery timeevery time
Whole test suite still green79%62%44%
First rule permanently broken atrule 47rule 30rule 24

The AI always did the job in front of it. All 1200 changes, in this run and in every other one. What it stopped doing was keeping the previous rules working.

And look where the middle column sits. We never mentioned a metric. We asked for tidy code in plain English. Retention still fell hard, and the first permanent breakage still arrived seventeen rules earlier than normal.

That answers the question we ran this for. Roughly half of the damage we blamed on the metric last time was not about the metric at all. It is the cost of restructuring while you are also trying to land a change. Telling the AI the formula then adds a second helping of damage on top, plus the static holders, plus the runs that disappear.

Why does restructuring cost correctness? We do not know for certain yet. The shape of it suggests the agent is doing two jobs in one turn. Land the new rule. Also rearrange the neighbourhood. Both get done. The blast radius of the second job is what breaks rule 19 while you are busy with rule 30.

What to take from this

If you are working with AI tools, here is the practical version.

Asking for structure works. One plain paragraph did most of what our exact scoring formula did. You do not need to invent a metric and feed it to the model. Plain words about single responsibility and putting logic where it belongs are enough to get most of the benefit.

Expect the bill to rise. Structure is work. The agent reads more and edits more. 23% more here. If someone tells you a prompt makes an agent write better code for free, be curious about what they measured.

A quiet metric is not a clean codebase. Our structural numbers all improved. Duplication went up, total code went up, files tripled, and the share of changes leaving the whole suite green fell from 79% to 62%. Only the tests caught that. Keep your acceptance tests, and run all of them, not just the ones for the thing you just changed.

Watch which way the refactoring goes. There is a real difference between "I moved this logic somewhere better" and "I put my new logic where the old code cannot charge me for it". The first one touches existing code. The second one avoids it. In these runs the plain prompt touched more existing code than normal, which is what real refactoring looks like. The formula prompt touched much less, which is what avoidance looks like.

Be suspicious of a pile of static helper classes. They are sometimes right. They are also the cheapest way to make almost any code metric look better, and in Spring they cost you injection, mocking, transactions and proxies.

The instruction has a target. On the Spring codebase, which had a real structural problem, the paragraph helped a lot. On the OfficeFloor codebase, which already spreads rules across wired functions, the same paragraph mostly bought churn. Some of its numbers got slightly worse. A structural instruction is a remedy. It is not a vitamin.

What we are not claiming

This is one paragraph of prompt, one model, one application and one sequence of sixty changes. A different wording would give different results. In particular the duplication finding is a result about our wording, and our wording is the one that asked for reuse.

We also cannot yet tell you how to get the structure without the correctness cost. That is the open problem.

For now the advice is short. Ask for good structure in words. Keep your scoring function to yourself. Budget for both.

Spending the Cost Function Without Telling It

Preprint / cs.SE / Empirical Software Engineering

Spending the Cost Function Without Telling It

A plain request for good structure. No formula. No gate. Most of the structural gain, on a bill that rises by 23 percent.

OfficeFloor · independent research · blog.officefloor.net

September 2026 · correspondence: daniel@officefloor.net


Abstract

The previous paper in this series handed an AI agent the structural cost function it was being scored by [9]. The agent optimised it, relocated the complexity, and its retention of previously passing rules fell from 0.787 to 0.440. That paper named one missing control. It could not separate disclosing a metric from any prompt that directs structural effort. This paper is that control. The agent is asked, in plain language, to put each piece of logic where it belongs and to keep units small. The metric is never mentioned.

Asking works, and most of the structural gain does not need the formula. Spring's per checkpoint change impact slope falls from 435.1 to 52.4. The controller file grows by 604 to 962 lines across the ten control chains. Across the ten cohesion chains it takes on 36 to 273. The heaviest class in the codebase stops being the controller in ten chains out of ten. Reachable complexity on the create path falls from 200.7 to 128.5, with every chain below the control's lowest.

The gain is bought, not gamed. Agent spend rises from $78.40 to $96.73 per Spring chain, an interval that excludes zero, and the agent takes about 320 more tool turns per chain. Complexity is relocated again, and this time the total goes up, from 282.3 to 325.6. Duplication rises from 2,233 to 2,818 clone lines even though the prompt asks for reuse. What does not happen is the escape. Container dispatched classes stay at 4.2 per chain against 12.7 under disclosure, and the worst chain reaches 8 against 33. The rules stay where the call graph can see them. New classes are idiomatic beans about as often as static utilities, 15.7 against 16.4, where disclosure produced 20.6 statics against 7.3 beans.

The correctness result is the reason the control was run. Retention of previously passing rules falls from 0.787 to 0.622 for Spring, against 0.440 under disclosure. The standing failure rate goes 0.99 percent, 1.76 percent, 2.86 percent across the three conditions. First permanent failure arrives at rule 47.5, then rule 30, then rule 24. So about half of the correctness cost previously charged to disclosure belongs instead to structural direction of any kind. Disclosure adds the rest, and adds the gaming. The narrow conclusion from the last paper survives and gains a price tag. Ask for structure in words. Expect to pay for it.

Keywords: change impact · prompt intervention · Goodhart's law · conservation of complexity · AI assisted development · software architecture · code degradation · agent cost

1Introduction

This series runs an AI agent through sixty accumulating change specifications on one REST endpoint, twice, on two architectures of the same application. The recurring Spring result is a handler that grows. The recurring rebuttal is that nobody told the agent to do better.

The previous paper told it, in the strongest form available [9]. It put the exact scoring arithmetic in the implement prompt at all sixty checkpoints. The agent solved the disclosed measure almost perfectly. It also relocated sixty business rules into twice as many files, wrote more duplication, moved two chains out of the call graph entirely, and stopped keeping earlier rules working from rule 24 onward.

That result had a hole in it, and the paper said so in its own threats section. The disclosed prompt was compared against a plain "implement it" control. So the correctness loss could belong to disclosure. It could equally belong to any instruction that makes the agent spend effort on structure while it is also trying to land a change. Those are very different findings. One says do not show the agent your metric. The other says restructuring under change pressure costs correctness whatever prompts it.

This paper reports the condition that separates them. The prompt asks for well placed, single responsibility code in ordinary words. It never names the metric, the formula, the experiment or the architecture. Everything else is held fixed. Whatever the plain request reproduces is not about disclosure.

There is a second question, and it turns out to be the more useful one for practitioners. The agent here is metered. Every checkpoint records dollars, tokens, tool turns and wall clock. So the cost of an instruction can be measured directly rather than assumed. The disclosed formula was free. This request is not.

  1. The load bearing control for the previous paper's correctness claim. A structure directed prompt that never mentions the metric, ungated, over 1200 checkpoints, against the same 1200 checkpoint control.
  2. An attribution of the correctness cost. Roughly half of the fall that disclosure produced is reproduced by plain words. The rest, and all of the gaming, needs the formula.
  3. A measured price for asking. Agent spend rises 23 percent for Spring and 17 percent for OfficeFloor, with tool turns and wall clock rising with it, on intervals that exclude zero.
  4. A second instance of conservation with the opposite sign. Under plain words the complexity moves and the total rises, where under the formula it moved and stayed flat.

2Background and related work

Goodhart's law needs no further demonstration after the last paper [5]. This one asks the question that follows it. If the measure cannot be the target, can words do the work instead, and what do the words cost.

Brooks separated the essential difficulty of a problem from the accidental difficulty of how we build it [3]. Tesler's conservation of complexity says design moves complexity rather than removing it [4]. The last paper found both, with Spring's total touched complexity flat while its distribution inverted. The interesting question for this condition is whether a request phrased in the vocabulary of good design behaves differently from a request phrased as arithmetic. It does behave differently. It is not obvious in advance which direction that difference runs.

The measures come from the same two sources as the rest of the series. SlopCodeBench supplies erosion, verbosity and degradation slope [1]. SWE-CI supplies normalized change, EvoScore and the zero regression rate [2]. Change impact itself is defined and externally validated against defects in human written repositories [6].

3The intervention

The control prompt is the instruction used unchanged throughout this series. Implement the specification. Make the tests pass. The cohesion prompt adds one paragraph:

Implement the following change to the application so that it fully
satisfies the specification and all existing and new tests pass.

As you do, keep the code well structured: put each piece of logic
where it belongs, in a small unit with a single responsibility, and
reuse existing code instead of copying it. Do not let any one class
or method grow into a catch-all that accumulates unrelated logic.
Run the test suite and make it green before finishing.

{spec}

What that paragraph does not contain matters as much as what it does. There is no formula or the existence of an experiment. It is the paragraph a senior engineer might add to a ticket. It asks for reuse explicitly, which becomes relevant in Section 5.2.

The comparison against the disclosed condition is therefore a comparison of two ways of asking for the same underlying property. One states the arithmetic. One states the intent.

4Study design

Research questions

  • RQ1. How much of the structural improvement survives when the metric is never disclosed?
  • RQ2. Is the improvement relocation again, and does the work stay visible to the measurement?
  • RQ3. What does the instruction cost in agent spend and time?
  • RQ4. How much of the correctness loss reported for disclosure belongs to disclosure, and how much to structural direction of any kind?

Arms, chains and checkpoints

The coding agent is held fixed. Architecture is the independent variable. One arm is a Spring @RestController codebase. The other is an OfficeFloor codebase of YAML composed functions. Both implement the same PetClinic REST application. Sixty change specifications land in sequence on the single endpoint POST /api/owners. Most are additive. Every fourth from the eighth is mutative, revising prior rules and shipping updated copies of the affected tests. There are ten independent chains per architecture, so this condition is 1200 agent turns. The run is blind-202609160027. The control is blind-202608100006. The disclosed condition is blind-202609010045. All three use the same specification file and the same acceptance suite, neither of which was touched between them, and all three ran the same model.

Blind protocol

The agent sees the current specification and the accumulated tests up to and including the current checkpoint. It never sees future ones. Each turn runs in a history less sandbox rebuilt from the worktree, so the agent cannot infer which checkpoint it is on from git history. Each turn gets a fresh configuration and memory. Acceptance tests are injected per checkpoint rather than pre committed. After the turn the full accumulated suite runs, and regressions are computed against the set of tests passing before the checkpoint.

Nothing gated this run

The gate is selected by strategy name and this strategy does not select it. Every one of the 1200 capture records carries an empty gate field and zero refactors. No change was discarded, no refactor turn ran, and all twenty chains reached checkpoint sixty. This matters for the comparison. The disclosed condition carried an active gate that fired on 3 of its 1200 checkpoints, which is why that paper reported itself as prompt only rather than purely so. This condition needs no such qualification. The prompt is the whole intervention.

Measures

Degradation slope is the OLS slope of a metric on checkpoint number, with 95 percent intervals from a bootstrap clustered on chains. Between arm tests are differences of those slopes under Benjamini Hochberg control over the whole metric family. Chain level comparisons between conditions, which this paper adds, resample the ten chains on each side and report the 95 percent interval of the difference of means. Alongside the scoped metrics there is an unscoped cumulative audit: one diff per chain from the branch base to its final commit, over every changed file, with each changed line attributed to the function containing it at the tip. Created classes are classified from the parsed function list. Duplication is measured by clone detection with a smell pass beside it.

One analyser, four runs

All runs in this series are pure derive. Structural metrics are recomputed from materialised worktrees at each checkpoint commit, so a metric added later applies to completed runs. Every number in this paper comes from a single re analysis pass over all four runs on 21 September 2026, with one analyser build and one tool set. That also closes a caveat from the previous paper, where the smell half of the verbosity metric had silently failed to run. It now runs for every capture. It contributes 26 to 27 lines per chain tip against 2,200 to 2,800 clone lines, so the duplication conclusions in that paper and this one rest on clones either way.

5Results

RQ1: most of the structure, none of the disclosure

Spring, slope per checkpointcontrolcohesion promptdisclosed formula
impact_composite435.152.49.71
impact_mutation241.938.03.17
impact_godclass193.214.46.54
wmc_handler1.7690.1840.038
wmc_max1.8080.8790.556
entry_cc0.12220.08660.035
erosion_handler0.003260.002440
node_cc_median3.0281.9851.101
erosion0.0015250.0008330.000127

Table 1. Per checkpoint OLS slopes over ten Spring chains, three conditions. wmc_handler is the weight of the class the endpoint routes through. entry_cc is the entry handler's own complexity. node_cc_median is the complexity reachable from one handling node. Intervals for the two columns compared here: impact_composite is 435.1 [307.2, 586.6] in the control and 52.4 [17.0, 96.6] under the cohesion prompt. wmc_handler is 1.769 [1.592, 1.925] and 0.184 [0.112, 0.265].

The plain request moves every structural slope in the same direction the formula did. Which intervention looks larger depends on how the comparison is framed, and both framings belong here. As a ratio the formula wins easily. Change impact falls by a factor of eight here against forty five there. The weight of the routed class falls by a factor of ten against forty seven. As an absolute quantity of decay removed, the plain request takes most of what was available. The control slope is 435.1. The cohesion prompt removes 383 of it and the formula removes 425. On wmc_handler the plain request removes 1.585 of the 1.731 the formula removed. The remaining difference between the two interventions is the tail, and Section 5.2 shows what the formula did to reach it. OfficeFloor moves too, from 76.4 to 32.0 on impact_composite, on a codebase that had much less to gain.

The plainest number is again not a slope. The controller file starts at 203 lines. Across the ten control chains it takes on 604 to 962 added lines. Across the ten cohesion chains it takes on 36 to 273. Under disclosure it took on 2 to 44. The growing handler that this series was built on is not eliminated here. It is cut to about a fifth of its size and it stops being the dominant object in the codebase. In nine of ten control chains the heaviest class at the tip is OwnerRestControllerV1. In ten of ten cohesion chains it is Owner, the entity, which is heavy because it holds accessors rather than decisions. Mean heaviest class weight falls from 136.6 to 75.9, an interval of [−75.0, −45.6].

Comprehension load on the create path falls with it. Summed reachable complexity from the declared entry node is 200.7 at the tip in the control and 128.5 under the cohesion prompt, and every one of the ten cohesion chains lands below the control's lowest chain. The median per checkpoint change impact falls from 4,520 to 1,388, an interval of [−4,020, −2,430].

The between arm test tells a more interesting story than the within arm slopes. Under the control prompt, Spring minus OfficeFloor is +1.769 [1.597, 1.929] on wmc_handler and +358.7 [234.2, 507.2] on impact_composite. Under the cohesion prompt the first becomes +0.184 [0.111, 0.261], still surviving FDR control, still with a Cliff's delta of +1.000, meaning every Spring chain still separates from every OfficeFloor chain. The second becomes +20.4 [−15.8, 65.3] and stops excluding zero. So a good prompt removes the change impact difference between the architectures while leaving the god class difference intact and perfectly separated, merely small. erosion_handler, the measure of complexity concentrating in the routed class, falls from +0.00326 [0.00160, 0.00489] to +0.00244 [0, 0.00493]. It stops surviving FDR control, but the point estimate only drops by a quarter. It loses significance through a wider interval, not through a vanished effect. We read that as weakened, not gone.

One between arm result runs against the series. node_cc_median, the per node comprehension load, still favours OfficeFloor at +1.93 [1.55, 2.26] with a large effect size. But node_path_cc, the summed complexity along the declared path, inverts: Spring's tip value is 128.5 against OfficeFloor's 201.1. That difference is recorded in the run's own counter signal table. Two things about it. First, unlike the disclosed condition, this inversion is not a measurement escape, and Section 5.2 gives the evidence. Second, the two arms are not like for like on this measure. OfficeFloor declares about twenty wired nodes and Spring declares one, so the path sum adds up twenty closures against one. The per node figure is the comparable one, and it still separates the arms.

RQ2: relocated again, and this time the total went up

Spring, base to tip, per chaincontrolcohesion promptdisclosed formula
CC sum over touched functions282.3 ± 22.2325.6 ± 19.4267.9 ± 41.5
  in files the run created46.3205.1218.2
  in pre existing files236.0120.549.7
distinct functions touched110.3178.5133.5
files parsed18.7 ± 5.458.0 ± 7.741.4 ± 3.3
production Java lines, final2,6992,8212,551
clone lines, final2,2332,8182,508
container dispatched classes, tip3.04.212.7

Table 2. The unscoped cumulative audit, plus the tip level counts that test for escape. Every changed line at the chain tip is attributed to the function that contains it, and a function counts once however many checkpoints edited it. No prompt side scoping can hide from this. Container dispatched classes are advice, aspect, filter and listener types, which the framework invokes and no call graph reaches. The application's own baseline is 3.

The complexity was relocated once more. Logic in pre existing files falls from 236.0 to 120.5. Logic in files the run created rises from 46.3 to 205.1. Files involved triple. That is Tesler's conservation again [4], and the sixty rules are Brooks's essential difficulty either way [3].

The sign of the total is the new part. Under the formula, Spring's total touched complexity was flat inside its spread, 282.3 against 267.9. Under plain words it rises to 325.6, and the chain level interval on the tip's whole codebase complexity excludes zero at [+26.5, +60.1]. Asking for good structure did not conserve complexity. It added some. The codebase is larger, not smaller: 2,821 production Java lines against 2,699 in the control, where the formula shrank it to 2,551. That is the cost of writing small units. Each one needs a declaration, a constructor, an injection point and a call site.

Duplication rose, which is the counter result of this paper. Clone lines go from 2,233 to 2,818, an interval of [+425, +732]. The prompt contains the sentence "reuse existing code instead of copying it". Duplication rose by 26 percent anyway, and rose further than it did under the formula, which never asked for reuse at all. We do not have a mechanism for this. The available guess is that dispersal into many small units makes a shared helper harder to find than to rewrite, and that the instruction to keep units small competes with the instruction to reuse. The honest statement is that the one explicit request in the paragraph is the one the run did not deliver.

Spring, classes created per chaincontrolcohesion promptdisclosed formula
total9.041.431.5
  injected bean0.015.77.3
  static utility2.316.420.6
  exception6.46.01.6
  instance class0.01.21.1

Table 3. What kind of class now holds a rule, classified from the parsed function list rather than a regex. Static utilities score well on the impact formula and give up injection, test seams, proxying and transaction participation. The cohesion run creates more classes than the disclosed run and makes about half of them beans.

This table is where the two interventions part company. Both disperse. They disperse into different things. Under the formula the dominant new object is the static utility, 20.6 per chain against 7.3 beans, because a static method in an empty class scores near zero on the term the formula punishes. Under plain words the split is 16.4 statics against 15.7 injected beans, on more created classes overall. A bean is the idiomatic Spring answer to "put this where it belongs". A static holder is the cheap answer to "minimise this product". The prompts got different code because they asked different questions, and only one of them was asking about the score.

Nothing left the measurement. Container dispatched classes stay at 4.2 per chain against the application's baseline of 3, with a worst chain of 8. Under disclosure that count was 12.7 with a worst chain of 33, and two chains had moved every rule into framework invoked interceptors where the call graph could not follow. Here the reduction in create path complexity is a reduction in create path complexity. That is why we are willing to report the node_path_cc inversion in Section 5.1 as a real measurement rather than an artifact.

One more number cuts against the tidy reading. Blast radius went up. The run modifies 209.8 pre existing functions per chain against 178.5 in the control, and the number of checkpoints that disturb nothing already there falls from 44 to 24 out of 600. Under disclosure both moved the other way, to 101.3 and 183. The plain request makes the agent go back into existing code and rearrange it. The formula made it avoid existing code, because existing code is what the formula charges for. Those are opposite behaviours, and only one of them is what a reviewer means by refactoring.

RQ3: the price of asking

per chain, 60 checkpointsSpring controlSpring cohesiondiff, 95% CIOF controlOF cohesion
agent spend, USD78.4096.73[+14.5, +22.1]86.51101.18
tool turns1,5221,844[+260, +384]1,8101,868
wall clock, hours4.044.62[+0.43, +0.72]5.095.22
output tokens, thousands625809739827
spend per checkpoint, USD1.311.611.441.69

Table 4. What the paragraph cost. Intervals are 95 percent bootstrap intervals on the difference of chain means, ten chains on each side. The OfficeFloor spend interval is [+10.7, +18.8]. For comparison, the disclosed formula cost nothing: $76.82 against $78.40 for Spring, on an interval of [−5.3, +1.9] that contains zero.

0.50 0.60 0.70 0.80 $75 $80 $85 $90 $95 $100 agent spend per chain (Spring, 60 checkpoints) retained correctness (strict_pass) control prompt impact slope 435.1 cohesion prompt impact slope 52.4 disclosed formula impact slope 9.7
Figure 1. Three prompts, Spring arm, ten chains each. Horizontal axis is what the agent spent. Vertical axis is the share of checkpoints leaving the whole accumulated suite green. The label under each point is that condition's per checkpoint change impact slope, which falls as the points move down. The best structure is at the bottom of the chart. The best retention is at the top, and it belongs to the prompt that produced the god method. Structure was bought with correctness in both interventions, and with money in only one of them.

The paragraph is not free. Spring spend rises 23 percent and OfficeFloor spend rises 17 percent, both on intervals that exclude zero. The agent takes about 320 more tool turns per Spring chain, roughly five more per change, and runs about half an hour longer. Output tokens rise by 29 percent. Across the full condition, twenty chains, the extra bill is about $330 on a base of about $1,650.

The disclosed formula bought a larger structural effect for nothing measurable. That is worth stating plainly, because it is the commercial argument for the thing this series advises against. Arithmetic is cheap to optimise. Judgement is not.

One caution on reading the money as a dial. Within the cohesion condition, chains that spent more did not finish better structured. The rank correlation between chain spend and tip create path complexity is +0.50 across ten chains, which is the wrong sign for a dial, and the correlation with median change impact is +0.22. So the 23 percent is the price of the instruction, not a knob that buys proportional structure. Paying more did not help. Being asked did.

RQ4: who owns the correctness cost

Springcontrolcohesion promptdisclosed formula
own rule delivered (func)1.0001.0001.000
whole suite green (strict_pass)0.7870.6220.440
standing prior failures, share0.99%1.76%2.86%
median chain onset of first failurerule 47.5rule 30rule 24
breakage on untouched rules373856
OfficeFloor, whole suite green0.7320.6450.577
OfficeFloor, standing failures1.03%1.82%1.85%

Table 5. Delivered correctness across the three conditions. Every checkpoint's own new rule landed in every chain of every condition, so func is 1.000 throughout and the degradation is entirely in retention. Standing failures count prior tests failing at a checkpoint, and a rule broken and never repaired keeps counting. Mutative checkpoints ship updated copies of the tests they revise, so these are real failures rather than intended churn. The Spring strict_pass fall against control is [−0.267, −0.062] and the standing failure rise is [+0.25, +1.34].

This is the result the condition was run for. A prompt that never mentions the metric still costs retention. Spring falls from 0.787 to 0.622 on an interval that excludes zero. Standing failures nearly double. The first permanent failure arrives seventeen rules earlier. None of that can be attributed to disclosure, because nothing was disclosed.

Take the disclosed condition's fall as the quantity to be explained. Spring's strict_pass dropped 0.347 from control to disclosure. The cohesion prompt reproduces 0.165 of it, which is 48 percent. On standing failure rate the rise is 1.86 points and the cohesion prompt reproduces 0.77, which is 41 percent. On onset the control is rule 47.5, the cohesion prompt is rule 30 and disclosure is rule 24. So the previous paper's headline correctness number is about half a disclosure effect and about half a restructuring effect. Both halves are real. Only one of them is a Goodhart problem.

Two qualifications keep this honest. First, breakage on untouched rules does not move at all under the cohesion prompt, 38 against the control's 37, where disclosure raised it to 56. That is a rare and concentrated event and the previous paper flagged it as the weaker of its correctness signals, but it points the same way: the plain prompt loses retention without the extra unintended breakage. Second, OfficeFloor's retention fall, 0.732 to 0.645, has an interval of [−0.183, +0.023] that contains zero. The standing failure rise for OfficeFloor does exclude zero. So the arm with less to restructure pays less, and on the headline measure its payment is not statistically distinguishable from noise.

The metric kept measuring

Change impact is not the objective in this condition, so its construct validity can be tested rather than assumed. Within the cohesion run it correlates with independently measured agent spend at Spearman +0.661 for Spring and +0.684 for OfficeFloor, with comprehension effort at +0.661 and +0.633, and with model time at +0.573 and +0.642. Spring's three are all higher than in the control run, which reads +0.536, +0.488 and +0.534. OfficeFloor's three sit within a few points of its control values of +0.686, +0.595 and +0.682, two of them slightly lower. Checkpoints that broke an untouched rule carry a median impact_composite of 12,290 against 1,256 for those that did not.

So the score still ranks changes by what they cost a maintainer, on a run that was pushed hard toward better structure by other means. A measure used as evidence keeps working. That is the half of the previous paper's conclusion this condition was able to test, and it survives.

6Discussion

What this control settles is narrow and it matters. The previous paper reported a large correctness loss under a disclosed cost function and could not say what caused it. About half of that loss now has a different owner. Ask an agent for good structure in ordinary words, with no metric anywhere near it, and retention still falls, standing failures still nearly double, and the first permanent break still arrives much earlier. Restructuring under a stream of accumulating change costs correctness. That is not a metric artifact. It looks like a property of doing two jobs in one turn.

What the control does not settle is the rest. Disclosure still costs a further 0.18 of retention beyond what plain words cost, and it brings behaviour plain words do not produce. Static utilities instead of beans. Two chains out of the call graph entirely. More unintended breakage. The dispersal under plain words looks like engineering. The dispersal under the formula looks like arbitrage against a specific term of a specific product.

For practice the useful finding is the price tag. A paragraph of ordinary structural instruction bought an eightfold reduction in change impact and a controller a fifth of the size, and it cost 23 percent more spend and about five more tool turns per change. That is a trade most teams would take. It should be made with open eyes on both sides of it. The same paragraph raised duplication by a quarter, raised total complexity, tripled the number of files, and cost a fifth of the suite's retention. It is not free, it is not purely positive, and the only instrument that reported the downside was the test suite.

For the series thesis the result is mixed, which is the right outcome for a control. The rebuttal that this series exists to test is that architecture does not matter because better prompting fixes it. Better prompting does fix a lot. The god method is cut to a fifth. The heaviest class stops being the controller in every chain. The change impact difference between the architectures stops excluding zero. And yet the per node comprehension load still separates the arms by +1.93 with a large effect size, the routed class difference still survives correction with every chain separated, and Spring paid 23 percent more spend to get there while OfficeFloor arrived at a similar distribution as its ordinary way of working. The honest summary is that a good prompt narrows the architectural gap substantially, pays money for the privilege, and does not close it.

One asymmetry is worth flagging for anyone applying this to their own codebase. OfficeFloor's median per checkpoint change impact went up under the cohesion prompt, from 319 to 472, on an interval that excludes zero, while its slope fell. Spring's fell hard on both. The prompt is worth the most where a structural problem exists. On a codebase whose structure is already imposed, the same instruction mostly buys churn. A structural instruction is a remedy, not a hygiene rule, and it has a target.

The tool guided refactor condition, in which a gate flags a change and one guided refactor runs without the agent ever seeing the metric, is also complete. It is reported separately, and a combined paper over all four conditions follows.

7Threats to validity

The prompt is one sample of a large space

This is one paragraph. Its four clauses, place logic where it belongs, keep units small and single purpose, reuse rather than copy, do not let a class become a catch all, could be reordered, softened or strengthened, and the result would move. The duplication finding in particular is a result about this wording, since the reuse clause is in it and reuse got worse. Nothing here measures the best achievable prompt. It measures a reasonable one.

Cost is vendor metered

Spend is what the agent's own accounting reports per turn, summed per chain. It is a faithful measure of what this run cost to execute. It is not a measure of engineering effort, and it is denominated in one vendor's pricing at one point in time. Tool turns and wall clock move with it, which is why all three are in Table 4 rather than the dollars alone.

Measurement

Cyclomatic complexity is a proxy for comprehension effort, not a measurement of it. The impact score charges only lines inside parsed function bodies, so declarative logic scores zero in both arms. node_path_cc sums closures over the declared path, and OfficeFloor declares about twenty nodes to Spring's one, so it must never be quoted as a like for like comparison. In the control run, three of ten Spring chains renamed the declared entry function, so their entry and path figures are blank and the control's 200.7 is a mean over seven chains rather than ten. The cohesion and disclosed runs resolve all ten. Duplication is clone detection plus a smell pass that contributes 26 to 27 lines against 2,200 to 2,800 clone lines, so it is effectively clone detection.

Run to run differences other than the prompt

The model is identical across all three conditions. The specification file and acceptance suite were last modified before the earliest of them. The harness moved between runs, and so did the agent CLI build, 2.1.222 for the control against 2.1.236 here. Analysis is identical by construction, since every number comes from one re derivation pass over all four runs. A CLI build difference is a real uncontrolled variable and it cannot be ruled out as a contributor to the spend difference in particular.

Statistical

Ten chains per architecture per condition. Slope intervals come from a bootstrap clustered on chains. Between arm verdicts are FDR controlled over the whole metric family, and the run publishes its own disagreements: 24 contradicted expectations and 18 predicted effects that did not appear, against 14 and 10 in the control and 34 and 20 under disclosure. The chain level comparisons between conditions are a difference of means over ten chains on each side, which is a small sample, and they are reported with intervals for that reason. Breakage on untouched rules remains a rare and concentrated event and should not be read as a trend in any condition.

Generality

One model. One endpoint. One sixty step checkpoint plan. Two codebases. The plan is fixed across conditions, which is what makes the comparison clean, and it also means the specific correctness numbers are properties of this plan. Nothing here establishes that the same paragraph would cost 23 percent on a different change stream.

8Reproducibility and data availability

Data availability All runs are reproducible from committed configuration. Each chain's branch carries its own configuration snapshot, its per checkpoint capture records, and a provenance manifest naming the model, the tool versions and the parser probes. Every number here is recomputed from the commits rather than trusted from a log. The harness, the checkpoint plan and the acceptance suite are open [7].
python -m harness.run_experiment --config config.yaml --test-mode blind \
       --strategy cohesion-prompt
python -m harness.analyze        --config config.yaml --run-id blind-202609160027

The control runs identically with the unchanged implement prompt. Analysis recomputes structural metrics from materialised worktrees at each checkpoint commit, so a metric added later can be applied to completed runs without re running the agent. The four runs in this paper were all re analysed in one pass on 21 September 2026.

9Conclusion

Asked in plain words to keep the code well structured, the agent did. Change impact fell eightfold. The controller that this series was built on grew by a fifth of what it grows under a neutral prompt, and stopped being the heaviest class in every chain. Reachable complexity on the create path fell by a third, and unlike the disclosed condition it fell where the measurement could still see it. The rules went into injected beans about as often as into static holders.

It was not free and it was not clean. Spend rose 23 percent. The complexity moved again and the total rose rather than held. Duplication rose by a quarter, in a run that was explicitly asked to reuse. And retention of earlier rules fell from 0.787 to 0.622.

That last number is the point of the experiment. The previous paper watched retention fall to 0.440 under a disclosed cost function and could not say whether the metric caused it. About half of the fall happens without any metric at all. Structural direction under a stream of change costs correctness on its own. Disclosure then adds a second helping, plus the static utilities, plus the chains that disappear from the call graph.

So the advice from the last paper stands and gets sharper. Keep the scoring function out of the agent's context and use it to watch. Ask for structure in words if you want structure, and budget for it. Then watch the test suite, because in this experiment it was the only instrument that noticed what the structural numbers were celebrating.


References and notes

  1. SlopCodeBench. No-context iterative extension, erosion and verbosity metrics, degradation slope. arXiv:2603.24755.
  2. SWE-CI. Continuous-integration gate, Normalized Change, EvoScore, Zero-Regression Rate. arXiv:2603.03823.
  3. F. P. Brooks Jr. No Silver Bullet: Essence and Accident in Software Engineering. Proc. IFIP Congress, 1986. Reprinted in IEEE Computer 20(4), 1987.
  4. L. Tesler. The law of conservation of complexity, articulated in the mid-1980s. Every process carries an irreducible complexity that design can move but not remove. Design decides only who bears it.
  5. C. A. E. Goodhart. Problems of Monetary Management: The U.K. Experience, 1975. The commonly quoted form is M. Strathern's restatement, 1997. When a measure becomes a target, it ceases to be a good measure.
  6. Change Impact in the Wild. An external test of the metric against defects in twenty human-written repositories. blog.officefloor.net.
  7. PetClinic-Evolve degradation study. Prior posts and harness. blog.officefloor.net.
  8. Cyclomatic complexity and per-function line ranges are computed with lizard. It is pinned and probed before each run. Clone detection uses jscpd at pinned thresholds. The smell pass is PMD against a ruleset that deliberately excludes every complexity rule.
  9. Telling the Agent the Cost Function. The disclosed-formula condition, run blind-202609010045, September 2026. blog.officefloor.net.

PREPRINT · OFFICEFLOOR · SEPTEMBER 2026

Wednesday, 16 September 2026

The decay you cannot see: what we measured and the one thing we will not claim

AI keeps the tests green while a codebase quietly rots. Here is how we measured that decay, what held up when we tried to break the metric, and why we refuse to sell it as a bug predictor.

The problem hiding behind green tests

AI is very good at complexity. That turns out to be the problem.

It keeps piling logic into code that already exists, because complexity does not slow it down the way it slows a person down. A god method grows another branch. A god class gains another method. The tests still pass. The feature still ships. Nobody notices.

Slowly the system decays to a point where no human can hold it in their head. Given enough of this, even the AI loses the thread, and every change starts breaking something else. By the time anyone feels the pain, the cheap fix is gone. What is left is an expensive, error prone refactor, or a rewrite.

The obvious question is whether you can see this coming. So we tried to measure it.

Score the change, not the code

Most quality metrics score a snapshot. They tell you how complex the code is right now. That is the wrong frame for decay, because decay is not a state, it is a motion. It is the act of adding to something already heavy.

So we measure the change, not the code. And we weight each change by how much existing tangled code it disturbs.

The core of the measure is the surrounding weight. That is the complexity already sitting in the class or module you are editing, measured on the state before your change. Adding a brand new file scores almost nothing, because nothing was there before. Growing an existing god class scores a lot. That asymmetry is the decay signal.

The full formula, and the three attempts it took to get there, are in an earlier post in this series. The short version is that a change costs more when it touches more files, disturbs more surrounding complexity, and adds more lines to already complex functions.

What the experiment showed

We ran two codebases through a long series of AI made changes. One was built additively, where you compose new pieces and wire them together. The other was built mutatively, where you edit a woven whole.

By the end of the run, a single change to the mutative codebase disturbed on the order of fifteen times more weighted structure than the same kind of change to the additive one. The decay was not hypothetical. It compounded, change after change, exactly as the thesis predicted.

Then we checked whether the number meant anything real, by correlating it against independent signals we captured during the run.

Signal Additive arm Mutative arm
Money cost of the change 0.69 0.54
Re-reading of existing code 0.60 0.49
Model time spent 0.68 0.53

The correlations held inside each arm independently, so this was not just an artifact of one codebase being harder overall. And when a change broke a rule in code it never touched, that change scored roughly seven to ten times higher than a change that broke nothing. The measure was tracking real cost, real comprehension load, and real blast damage.

This is the part where most tool posts would stop.

Then we tried to break our own metric

A measure that only ever flatters itself is worthless. So we took it to a wider, messier world. We scanned twenty open source repositories, across eight well sampled languages, and collected 349,165 per commit observations of the measure in the wild.

Then we asked the uncomfortable question. If this number is good, it should help predict where the bugs are. Does it?

It does not. Across that corpus, the measure does not beat plain file size as a defect predictor. File size is a famously strong and famously simple baseline for bug proneness, and our carefully weighted change measure did not clear it.

We could have quietly not run that test. We ran it, and we are telling you the result, because the result matters.

Why that is the right outcome

A negative result is only a failure if you were claiming the thing it disproves. We were not.

The measure was never built to predict defects. It was built to measure change effort and structural decay. Those are different quantities. File size predicts bugs well precisely because bigger files have more surface for anything to go wrong. That says little about how much a specific change grows the parts of your codebase that are becoming unmaintainable.

So we do not sell a bug predictor. We would be competing with file size and losing, and we would be lying. We measure the thing the experiment actually validated: how much effort a change costs, and how much it adds to the decay. That is a different and, we would argue, more actionable thing. It tells you where change is getting expensive, while it is still cheap to fix.

What it is actually good for

Three things.

It gates decay. When a change piles too much onto already heavy code, it can warn, or block the merge, and ask you to simplify or refactor first.

It routes attention. Every score comes with a ranked list of the files where the decay is concentrating, so a class quietly growing into a god class surfaces as a refactor candidate before it blocks anything.

And it grades on a curve, because a raw threshold cannot work. Change impact varies by around 300 times across projects and 60 times across languages. So instead of guessing a number, it grades a change by its percentile against a real distribution, blending a seed corpus with your own repository history.

The tool

All of this ships as ImpactGate. It runs as a command line tool, a git pre-commit hook, and as a check on GitHub, GitLab, and Jenkins. It treats an AI written change like any other change. It scores the blast radius, and the big ones stop for a human to look at while the fix is still small.

You can try it out with the following on your current code change:

docker run --rm -v "$PWD:/repo" ghcr.io/officefloor/impact-gate score

Honesty is the point

We think the interesting story here is not that we built a metric that tells a nice tale. It is easy to invent one of those. The interesting story is that we took our own metric to a place where it could fail, watched it fail at a job it was never meant to do, and kept the job it is good at.

If you are going to let AI write a lot of your code, you want tools that are clear about exactly what they measure, and honest about what they do not. This is a tool to highlight decay early, before it becomes expensive to fix.