Tuesday, 22 September 2026

Spending the Cost Function Without Telling It

Preprint / cs.SE / Empirical Software Engineering

Spending the Cost Function Without Telling It

A plain request for good structure. No formula. No gate. Most of the structural gain, on a bill that rises by 23 percent.

OfficeFloor · independent research · blog.officefloor.net

September 2026 · correspondence: daniel@officefloor.net


Abstract

The previous paper in this series handed an AI agent the structural cost function it was being scored by [9]. The agent optimised it, relocated the complexity, and its retention of previously passing rules fell from 0.787 to 0.440. That paper named one missing control. It could not separate disclosing a metric from any prompt that directs structural effort. This paper is that control. The agent is asked, in plain language, to put each piece of logic where it belongs and to keep units small. The metric is never mentioned.

Asking works, and most of the structural gain does not need the formula. Spring's per checkpoint change impact slope falls from 435.1 to 52.4. The controller file grows by 604 to 962 lines across the ten control chains. Across the ten cohesion chains it takes on 36 to 273. The heaviest class in the codebase stops being the controller in ten chains out of ten. Reachable complexity on the create path falls from 200.7 to 128.5, with every chain below the control's lowest.

The gain is bought, not gamed. Agent spend rises from $78.40 to $96.73 per Spring chain, an interval that excludes zero, and the agent takes about 320 more tool turns per chain. Complexity is relocated again, and this time the total goes up, from 282.3 to 325.6. Duplication rises from 2,233 to 2,818 clone lines even though the prompt asks for reuse. What does not happen is the escape. Container dispatched classes stay at 4.2 per chain against 12.7 under disclosure, and the worst chain reaches 8 against 33. The rules stay where the call graph can see them. New classes are idiomatic beans about as often as static utilities, 15.7 against 16.4, where disclosure produced 20.6 statics against 7.3 beans.

The correctness result is the reason the control was run. Retention of previously passing rules falls from 0.787 to 0.622 for Spring, against 0.440 under disclosure. The standing failure rate goes 0.99 percent, 1.76 percent, 2.86 percent across the three conditions. First permanent failure arrives at rule 47.5, then rule 30, then rule 24. So about half of the correctness cost previously charged to disclosure belongs instead to structural direction of any kind. Disclosure adds the rest, and adds the gaming. The narrow conclusion from the last paper survives and gains a price tag. Ask for structure in words. Expect to pay for it.

Keywords: change impact · prompt intervention · Goodhart's law · conservation of complexity · AI assisted development · software architecture · code degradation · agent cost

1Introduction

This series runs an AI agent through sixty accumulating change specifications on one REST endpoint, twice, on two architectures of the same application. The recurring Spring result is a handler that grows. The recurring rebuttal is that nobody told the agent to do better.

The previous paper told it, in the strongest form available [9]. It put the exact scoring arithmetic in the implement prompt at all sixty checkpoints. The agent solved the disclosed measure almost perfectly. It also relocated sixty business rules into twice as many files, wrote more duplication, moved two chains out of the call graph entirely, and stopped keeping earlier rules working from rule 24 onward.

That result had a hole in it, and the paper said so in its own threats section. The disclosed prompt was compared against a plain "implement it" control. So the correctness loss could belong to disclosure. It could equally belong to any instruction that makes the agent spend effort on structure while it is also trying to land a change. Those are very different findings. One says do not show the agent your metric. The other says restructuring under change pressure costs correctness whatever prompts it.

This paper reports the condition that separates them. The prompt asks for well placed, single responsibility code in ordinary words. It never names the metric, the formula, the experiment or the architecture. Everything else is held fixed. Whatever the plain request reproduces is not about disclosure.

There is a second question, and it turns out to be the more useful one for practitioners. The agent here is metered. Every checkpoint records dollars, tokens, tool turns and wall clock. So the cost of an instruction can be measured directly rather than assumed. The disclosed formula was free. This request is not.

  1. The load bearing control for the previous paper's correctness claim. A structure directed prompt that never mentions the metric, ungated, over 1200 checkpoints, against the same 1200 checkpoint control.
  2. An attribution of the correctness cost. Roughly half of the fall that disclosure produced is reproduced by plain words. The rest, and all of the gaming, needs the formula.
  3. A measured price for asking. Agent spend rises 23 percent for Spring and 17 percent for OfficeFloor, with tool turns and wall clock rising with it, on intervals that exclude zero.
  4. A second instance of conservation with the opposite sign. Under plain words the complexity moves and the total rises, where under the formula it moved and stayed flat.

2Background and related work

Goodhart's law needs no further demonstration after the last paper [5]. This one asks the question that follows it. If the measure cannot be the target, can words do the work instead, and what do the words cost.

Brooks separated the essential difficulty of a problem from the accidental difficulty of how we build it [3]. Tesler's conservation of complexity says design moves complexity rather than removing it [4]. The last paper found both, with Spring's total touched complexity flat while its distribution inverted. The interesting question for this condition is whether a request phrased in the vocabulary of good design behaves differently from a request phrased as arithmetic. It does behave differently. It is not obvious in advance which direction that difference runs.

The measures come from the same two sources as the rest of the series. SlopCodeBench supplies erosion, verbosity and degradation slope [1]. SWE-CI supplies normalized change, EvoScore and the zero regression rate [2]. Change impact itself is defined and externally validated against defects in human written repositories [6].

3The intervention

The control prompt is the instruction used unchanged throughout this series. Implement the specification. Make the tests pass. The cohesion prompt adds one paragraph:

Implement the following change to the application so that it fully
satisfies the specification and all existing and new tests pass.

As you do, keep the code well structured: put each piece of logic
where it belongs, in a small unit with a single responsibility, and
reuse existing code instead of copying it. Do not let any one class
or method grow into a catch-all that accumulates unrelated logic.
Run the test suite and make it green before finishing.

{spec}

What that paragraph does not contain matters as much as what it does. There is no formula or the existence of an experiment. It is the paragraph a senior engineer might add to a ticket. It asks for reuse explicitly, which becomes relevant in Section 5.2.

The comparison against the disclosed condition is therefore a comparison of two ways of asking for the same underlying property. One states the arithmetic. One states the intent.

4Study design

Research questions

  • RQ1. How much of the structural improvement survives when the metric is never disclosed?
  • RQ2. Is the improvement relocation again, and does the work stay visible to the measurement?
  • RQ3. What does the instruction cost in agent spend and time?
  • RQ4. How much of the correctness loss reported for disclosure belongs to disclosure, and how much to structural direction of any kind?

Arms, chains and checkpoints

The coding agent is held fixed. Architecture is the independent variable. One arm is a Spring @RestController codebase. The other is an OfficeFloor codebase of YAML composed functions. Both implement the same PetClinic REST application. Sixty change specifications land in sequence on the single endpoint POST /api/owners. Most are additive. Every fourth from the eighth is mutative, revising prior rules and shipping updated copies of the affected tests. There are ten independent chains per architecture, so this condition is 1200 agent turns. The run is blind-202609160027. The control is blind-202608100006. The disclosed condition is blind-202609010045. All three use the same specification file and the same acceptance suite, neither of which was touched between them, and all three ran the same model.

Blind protocol

The agent sees the current specification and the accumulated tests up to and including the current checkpoint. It never sees future ones. Each turn runs in a history less sandbox rebuilt from the worktree, so the agent cannot infer which checkpoint it is on from git history. Each turn gets a fresh configuration and memory. Acceptance tests are injected per checkpoint rather than pre committed. After the turn the full accumulated suite runs, and regressions are computed against the set of tests passing before the checkpoint.

Nothing gated this run

The gate is selected by strategy name and this strategy does not select it. Every one of the 1200 capture records carries an empty gate field and zero refactors. No change was discarded, no refactor turn ran, and all twenty chains reached checkpoint sixty. This matters for the comparison. The disclosed condition carried an active gate that fired on 3 of its 1200 checkpoints, which is why that paper reported itself as prompt only rather than purely so. This condition needs no such qualification. The prompt is the whole intervention.

Measures

Degradation slope is the OLS slope of a metric on checkpoint number, with 95 percent intervals from a bootstrap clustered on chains. Between arm tests are differences of those slopes under Benjamini Hochberg control over the whole metric family. Chain level comparisons between conditions, which this paper adds, resample the ten chains on each side and report the 95 percent interval of the difference of means. Alongside the scoped metrics there is an unscoped cumulative audit: one diff per chain from the branch base to its final commit, over every changed file, with each changed line attributed to the function containing it at the tip. Created classes are classified from the parsed function list. Duplication is measured by clone detection with a smell pass beside it.

One analyser, four runs

All runs in this series are pure derive. Structural metrics are recomputed from materialised worktrees at each checkpoint commit, so a metric added later applies to completed runs. Every number in this paper comes from a single re analysis pass over all four runs on 21 September 2026, with one analyser build and one tool set. That also closes a caveat from the previous paper, where the smell half of the verbosity metric had silently failed to run. It now runs for every capture. It contributes 26 to 27 lines per chain tip against 2,200 to 2,800 clone lines, so the duplication conclusions in that paper and this one rest on clones either way.

5Results

RQ1: most of the structure, none of the disclosure

Spring, slope per checkpointcontrolcohesion promptdisclosed formula
impact_composite435.152.49.71
impact_mutation241.938.03.17
impact_godclass193.214.46.54
wmc_handler1.7690.1840.038
wmc_max1.8080.8790.556
entry_cc0.12220.08660.035
erosion_handler0.003260.002440
node_cc_median3.0281.9851.101
erosion0.0015250.0008330.000127

Table 1. Per checkpoint OLS slopes over ten Spring chains, three conditions. wmc_handler is the weight of the class the endpoint routes through. entry_cc is the entry handler's own complexity. node_cc_median is the complexity reachable from one handling node. Intervals for the two columns compared here: impact_composite is 435.1 [307.2, 586.6] in the control and 52.4 [17.0, 96.6] under the cohesion prompt. wmc_handler is 1.769 [1.592, 1.925] and 0.184 [0.112, 0.265].

The plain request moves every structural slope in the same direction the formula did. Which intervention looks larger depends on how the comparison is framed, and both framings belong here. As a ratio the formula wins easily. Change impact falls by a factor of eight here against forty five there. The weight of the routed class falls by a factor of ten against forty seven. As an absolute quantity of decay removed, the plain request takes most of what was available. The control slope is 435.1. The cohesion prompt removes 383 of it and the formula removes 425. On wmc_handler the plain request removes 1.585 of the 1.731 the formula removed. The remaining difference between the two interventions is the tail, and Section 5.2 shows what the formula did to reach it. OfficeFloor moves too, from 76.4 to 32.0 on impact_composite, on a codebase that had much less to gain.

The plainest number is again not a slope. The controller file starts at 203 lines. Across the ten control chains it takes on 604 to 962 added lines. Across the ten cohesion chains it takes on 36 to 273. Under disclosure it took on 2 to 44. The growing handler that this series was built on is not eliminated here. It is cut to about a fifth of its size and it stops being the dominant object in the codebase. In nine of ten control chains the heaviest class at the tip is OwnerRestControllerV1. In ten of ten cohesion chains it is Owner, the entity, which is heavy because it holds accessors rather than decisions. Mean heaviest class weight falls from 136.6 to 75.9, an interval of [−75.0, −45.6].

Comprehension load on the create path falls with it. Summed reachable complexity from the declared entry node is 200.7 at the tip in the control and 128.5 under the cohesion prompt, and every one of the ten cohesion chains lands below the control's lowest chain. The median per checkpoint change impact falls from 4,520 to 1,388, an interval of [−4,020, −2,430].

The between arm test tells a more interesting story than the within arm slopes. Under the control prompt, Spring minus OfficeFloor is +1.769 [1.597, 1.929] on wmc_handler and +358.7 [234.2, 507.2] on impact_composite. Under the cohesion prompt the first becomes +0.184 [0.111, 0.261], still surviving FDR control, still with a Cliff's delta of +1.000, meaning every Spring chain still separates from every OfficeFloor chain. The second becomes +20.4 [−15.8, 65.3] and stops excluding zero. So a good prompt removes the change impact difference between the architectures while leaving the god class difference intact and perfectly separated, merely small. erosion_handler, the measure of complexity concentrating in the routed class, falls from +0.00326 [0.00160, 0.00489] to +0.00244 [0, 0.00493]. It stops surviving FDR control, but the point estimate only drops by a quarter. It loses significance through a wider interval, not through a vanished effect. We read that as weakened, not gone.

One between arm result runs against the series. node_cc_median, the per node comprehension load, still favours OfficeFloor at +1.93 [1.55, 2.26] with a large effect size. But node_path_cc, the summed complexity along the declared path, inverts: Spring's tip value is 128.5 against OfficeFloor's 201.1. That difference is recorded in the run's own counter signal table. Two things about it. First, unlike the disclosed condition, this inversion is not a measurement escape, and Section 5.2 gives the evidence. Second, the two arms are not like for like on this measure. OfficeFloor declares about twenty wired nodes and Spring declares one, so the path sum adds up twenty closures against one. The per node figure is the comparable one, and it still separates the arms.

RQ2: relocated again, and this time the total went up

Spring, base to tip, per chaincontrolcohesion promptdisclosed formula
CC sum over touched functions282.3 ± 22.2325.6 ± 19.4267.9 ± 41.5
  in files the run created46.3205.1218.2
  in pre existing files236.0120.549.7
distinct functions touched110.3178.5133.5
files parsed18.7 ± 5.458.0 ± 7.741.4 ± 3.3
production Java lines, final2,6992,8212,551
clone lines, final2,2332,8182,508
container dispatched classes, tip3.04.212.7

Table 2. The unscoped cumulative audit, plus the tip level counts that test for escape. Every changed line at the chain tip is attributed to the function that contains it, and a function counts once however many checkpoints edited it. No prompt side scoping can hide from this. Container dispatched classes are advice, aspect, filter and listener types, which the framework invokes and no call graph reaches. The application's own baseline is 3.

The complexity was relocated once more. Logic in pre existing files falls from 236.0 to 120.5. Logic in files the run created rises from 46.3 to 205.1. Files involved triple. That is Tesler's conservation again [4], and the sixty rules are Brooks's essential difficulty either way [3].

The sign of the total is the new part. Under the formula, Spring's total touched complexity was flat inside its spread, 282.3 against 267.9. Under plain words it rises to 325.6, and the chain level interval on the tip's whole codebase complexity excludes zero at [+26.5, +60.1]. Asking for good structure did not conserve complexity. It added some. The codebase is larger, not smaller: 2,821 production Java lines against 2,699 in the control, where the formula shrank it to 2,551. That is the cost of writing small units. Each one needs a declaration, a constructor, an injection point and a call site.

Duplication rose, which is the counter result of this paper. Clone lines go from 2,233 to 2,818, an interval of [+425, +732]. The prompt contains the sentence "reuse existing code instead of copying it". Duplication rose by 26 percent anyway, and rose further than it did under the formula, which never asked for reuse at all. We do not have a mechanism for this. The available guess is that dispersal into many small units makes a shared helper harder to find than to rewrite, and that the instruction to keep units small competes with the instruction to reuse. The honest statement is that the one explicit request in the paragraph is the one the run did not deliver.

Spring, classes created per chaincontrolcohesion promptdisclosed formula
total9.041.431.5
  injected bean0.015.77.3
  static utility2.316.420.6
  exception6.46.01.6
  instance class0.01.21.1

Table 3. What kind of class now holds a rule, classified from the parsed function list rather than a regex. Static utilities score well on the impact formula and give up injection, test seams, proxying and transaction participation. The cohesion run creates more classes than the disclosed run and makes about half of them beans.

This table is where the two interventions part company. Both disperse. They disperse into different things. Under the formula the dominant new object is the static utility, 20.6 per chain against 7.3 beans, because a static method in an empty class scores near zero on the term the formula punishes. Under plain words the split is 16.4 statics against 15.7 injected beans, on more created classes overall. A bean is the idiomatic Spring answer to "put this where it belongs". A static holder is the cheap answer to "minimise this product". The prompts got different code because they asked different questions, and only one of them was asking about the score.

Nothing left the measurement. Container dispatched classes stay at 4.2 per chain against the application's baseline of 3, with a worst chain of 8. Under disclosure that count was 12.7 with a worst chain of 33, and two chains had moved every rule into framework invoked interceptors where the call graph could not follow. Here the reduction in create path complexity is a reduction in create path complexity. That is why we are willing to report the node_path_cc inversion in Section 5.1 as a real measurement rather than an artifact.

One more number cuts against the tidy reading. Blast radius went up. The run modifies 209.8 pre existing functions per chain against 178.5 in the control, and the number of checkpoints that disturb nothing already there falls from 44 to 24 out of 600. Under disclosure both moved the other way, to 101.3 and 183. The plain request makes the agent go back into existing code and rearrange it. The formula made it avoid existing code, because existing code is what the formula charges for. Those are opposite behaviours, and only one of them is what a reviewer means by refactoring.

RQ3: the price of asking

per chain, 60 checkpointsSpring controlSpring cohesiondiff, 95% CIOF controlOF cohesion
agent spend, USD78.4096.73[+14.5, +22.1]86.51101.18
tool turns1,5221,844[+260, +384]1,8101,868
wall clock, hours4.044.62[+0.43, +0.72]5.095.22
output tokens, thousands625809739827
spend per checkpoint, USD1.311.611.441.69

Table 4. What the paragraph cost. Intervals are 95 percent bootstrap intervals on the difference of chain means, ten chains on each side. The OfficeFloor spend interval is [+10.7, +18.8]. For comparison, the disclosed formula cost nothing: $76.82 against $78.40 for Spring, on an interval of [−5.3, +1.9] that contains zero.

0.50 0.60 0.70 0.80 $75 $80 $85 $90 $95 $100 agent spend per chain (Spring, 60 checkpoints) retained correctness (strict_pass) control prompt impact slope 435.1 cohesion prompt impact slope 52.4 disclosed formula impact slope 9.7
Figure 1. Three prompts, Spring arm, ten chains each. Horizontal axis is what the agent spent. Vertical axis is the share of checkpoints leaving the whole accumulated suite green. The label under each point is that condition's per checkpoint change impact slope, which falls as the points move down. The best structure is at the bottom of the chart. The best retention is at the top, and it belongs to the prompt that produced the god method. Structure was bought with correctness in both interventions, and with money in only one of them.

The paragraph is not free. Spring spend rises 23 percent and OfficeFloor spend rises 17 percent, both on intervals that exclude zero. The agent takes about 320 more tool turns per Spring chain, roughly five more per change, and runs about half an hour longer. Output tokens rise by 29 percent. Across the full condition, twenty chains, the extra bill is about $330 on a base of about $1,650.

The disclosed formula bought a larger structural effect for nothing measurable. That is worth stating plainly, because it is the commercial argument for the thing this series advises against. Arithmetic is cheap to optimise. Judgement is not.

One caution on reading the money as a dial. Within the cohesion condition, chains that spent more did not finish better structured. The rank correlation between chain spend and tip create path complexity is +0.50 across ten chains, which is the wrong sign for a dial, and the correlation with median change impact is +0.22. So the 23 percent is the price of the instruction, not a knob that buys proportional structure. Paying more did not help. Being asked did.

RQ4: who owns the correctness cost

Springcontrolcohesion promptdisclosed formula
own rule delivered (func)1.0001.0001.000
whole suite green (strict_pass)0.7870.6220.440
standing prior failures, share0.99%1.76%2.86%
median chain onset of first failurerule 47.5rule 30rule 24
breakage on untouched rules373856
OfficeFloor, whole suite green0.7320.6450.577
OfficeFloor, standing failures1.03%1.82%1.85%

Table 5. Delivered correctness across the three conditions. Every checkpoint's own new rule landed in every chain of every condition, so func is 1.000 throughout and the degradation is entirely in retention. Standing failures count prior tests failing at a checkpoint, and a rule broken and never repaired keeps counting. Mutative checkpoints ship updated copies of the tests they revise, so these are real failures rather than intended churn. The Spring strict_pass fall against control is [−0.267, −0.062] and the standing failure rise is [+0.25, +1.34].

This is the result the condition was run for. A prompt that never mentions the metric still costs retention. Spring falls from 0.787 to 0.622 on an interval that excludes zero. Standing failures nearly double. The first permanent failure arrives seventeen rules earlier. None of that can be attributed to disclosure, because nothing was disclosed.

Take the disclosed condition's fall as the quantity to be explained. Spring's strict_pass dropped 0.347 from control to disclosure. The cohesion prompt reproduces 0.165 of it, which is 48 percent. On standing failure rate the rise is 1.86 points and the cohesion prompt reproduces 0.77, which is 41 percent. On onset the control is rule 47.5, the cohesion prompt is rule 30 and disclosure is rule 24. So the previous paper's headline correctness number is about half a disclosure effect and about half a restructuring effect. Both halves are real. Only one of them is a Goodhart problem.

Two qualifications keep this honest. First, breakage on untouched rules does not move at all under the cohesion prompt, 38 against the control's 37, where disclosure raised it to 56. That is a rare and concentrated event and the previous paper flagged it as the weaker of its correctness signals, but it points the same way: the plain prompt loses retention without the extra unintended breakage. Second, OfficeFloor's retention fall, 0.732 to 0.645, has an interval of [−0.183, +0.023] that contains zero. The standing failure rise for OfficeFloor does exclude zero. So the arm with less to restructure pays less, and on the headline measure its payment is not statistically distinguishable from noise.

The metric kept measuring

Change impact is not the objective in this condition, so its construct validity can be tested rather than assumed. Within the cohesion run it correlates with independently measured agent spend at Spearman +0.661 for Spring and +0.684 for OfficeFloor, with comprehension effort at +0.661 and +0.633, and with model time at +0.573 and +0.642. Spring's three are all higher than in the control run, which reads +0.536, +0.488 and +0.534. OfficeFloor's three sit within a few points of its control values of +0.686, +0.595 and +0.682, two of them slightly lower. Checkpoints that broke an untouched rule carry a median impact_composite of 12,290 against 1,256 for those that did not.

So the score still ranks changes by what they cost a maintainer, on a run that was pushed hard toward better structure by other means. A measure used as evidence keeps working. That is the half of the previous paper's conclusion this condition was able to test, and it survives.

6Discussion

What this control settles is narrow and it matters. The previous paper reported a large correctness loss under a disclosed cost function and could not say what caused it. About half of that loss now has a different owner. Ask an agent for good structure in ordinary words, with no metric anywhere near it, and retention still falls, standing failures still nearly double, and the first permanent break still arrives much earlier. Restructuring under a stream of accumulating change costs correctness. That is not a metric artifact. It looks like a property of doing two jobs in one turn.

What the control does not settle is the rest. Disclosure still costs a further 0.18 of retention beyond what plain words cost, and it brings behaviour plain words do not produce. Static utilities instead of beans. Two chains out of the call graph entirely. More unintended breakage. The dispersal under plain words looks like engineering. The dispersal under the formula looks like arbitrage against a specific term of a specific product.

For practice the useful finding is the price tag. A paragraph of ordinary structural instruction bought an eightfold reduction in change impact and a controller a fifth of the size, and it cost 23 percent more spend and about five more tool turns per change. That is a trade most teams would take. It should be made with open eyes on both sides of it. The same paragraph raised duplication by a quarter, raised total complexity, tripled the number of files, and cost a fifth of the suite's retention. It is not free, it is not purely positive, and the only instrument that reported the downside was the test suite.

For the series thesis the result is mixed, which is the right outcome for a control. The rebuttal that this series exists to test is that architecture does not matter because better prompting fixes it. Better prompting does fix a lot. The god method is cut to a fifth. The heaviest class stops being the controller in every chain. The change impact difference between the architectures stops excluding zero. And yet the per node comprehension load still separates the arms by +1.93 with a large effect size, the routed class difference still survives correction with every chain separated, and Spring paid 23 percent more spend to get there while OfficeFloor arrived at a similar distribution as its ordinary way of working. The honest summary is that a good prompt narrows the architectural gap substantially, pays money for the privilege, and does not close it.

One asymmetry is worth flagging for anyone applying this to their own codebase. OfficeFloor's median per checkpoint change impact went up under the cohesion prompt, from 319 to 472, on an interval that excludes zero, while its slope fell. Spring's fell hard on both. The prompt is worth the most where a structural problem exists. On a codebase whose structure is already imposed, the same instruction mostly buys churn. A structural instruction is a remedy, not a hygiene rule, and it has a target.

The tool guided refactor condition, in which a gate flags a change and one guided refactor runs without the agent ever seeing the metric, is also complete. It is reported separately, and a combined paper over all four conditions follows.

7Threats to validity

The prompt is one sample of a large space

This is one paragraph. Its four clauses, place logic where it belongs, keep units small and single purpose, reuse rather than copy, do not let a class become a catch all, could be reordered, softened or strengthened, and the result would move. The duplication finding in particular is a result about this wording, since the reuse clause is in it and reuse got worse. Nothing here measures the best achievable prompt. It measures a reasonable one.

Cost is vendor metered

Spend is what the agent's own accounting reports per turn, summed per chain. It is a faithful measure of what this run cost to execute. It is not a measure of engineering effort, and it is denominated in one vendor's pricing at one point in time. Tool turns and wall clock move with it, which is why all three are in Table 4 rather than the dollars alone.

Measurement

Cyclomatic complexity is a proxy for comprehension effort, not a measurement of it. The impact score charges only lines inside parsed function bodies, so declarative logic scores zero in both arms. node_path_cc sums closures over the declared path, and OfficeFloor declares about twenty nodes to Spring's one, so it must never be quoted as a like for like comparison. In the control run, three of ten Spring chains renamed the declared entry function, so their entry and path figures are blank and the control's 200.7 is a mean over seven chains rather than ten. The cohesion and disclosed runs resolve all ten. Duplication is clone detection plus a smell pass that contributes 26 to 27 lines against 2,200 to 2,800 clone lines, so it is effectively clone detection.

Run to run differences other than the prompt

The model is identical across all three conditions. The specification file and acceptance suite were last modified before the earliest of them. The harness moved between runs, and so did the agent CLI build, 2.1.222 for the control against 2.1.236 here. Analysis is identical by construction, since every number comes from one re derivation pass over all four runs. A CLI build difference is a real uncontrolled variable and it cannot be ruled out as a contributor to the spend difference in particular.

Statistical

Ten chains per architecture per condition. Slope intervals come from a bootstrap clustered on chains. Between arm verdicts are FDR controlled over the whole metric family, and the run publishes its own disagreements: 24 contradicted expectations and 18 predicted effects that did not appear, against 14 and 10 in the control and 34 and 20 under disclosure. The chain level comparisons between conditions are a difference of means over ten chains on each side, which is a small sample, and they are reported with intervals for that reason. Breakage on untouched rules remains a rare and concentrated event and should not be read as a trend in any condition.

Generality

One model. One endpoint. One sixty step checkpoint plan. Two codebases. The plan is fixed across conditions, which is what makes the comparison clean, and it also means the specific correctness numbers are properties of this plan. Nothing here establishes that the same paragraph would cost 23 percent on a different change stream.

8Reproducibility and data availability

Data availability All runs are reproducible from committed configuration. Each chain's branch carries its own configuration snapshot, its per checkpoint capture records, and a provenance manifest naming the model, the tool versions and the parser probes. Every number here is recomputed from the commits rather than trusted from a log. The harness, the checkpoint plan and the acceptance suite are open [7].
python -m harness.run_experiment --config config.yaml --test-mode blind \
       --strategy cohesion-prompt
python -m harness.analyze        --config config.yaml --run-id blind-202609160027

The control runs identically with the unchanged implement prompt. Analysis recomputes structural metrics from materialised worktrees at each checkpoint commit, so a metric added later can be applied to completed runs without re running the agent. The four runs in this paper were all re analysed in one pass on 21 September 2026.

9Conclusion

Asked in plain words to keep the code well structured, the agent did. Change impact fell eightfold. The controller that this series was built on grew by a fifth of what it grows under a neutral prompt, and stopped being the heaviest class in every chain. Reachable complexity on the create path fell by a third, and unlike the disclosed condition it fell where the measurement could still see it. The rules went into injected beans about as often as into static holders.

It was not free and it was not clean. Spend rose 23 percent. The complexity moved again and the total rose rather than held. Duplication rose by a quarter, in a run that was explicitly asked to reuse. And retention of earlier rules fell from 0.787 to 0.622.

That last number is the point of the experiment. The previous paper watched retention fall to 0.440 under a disclosed cost function and could not say whether the metric caused it. About half of the fall happens without any metric at all. Structural direction under a stream of change costs correctness on its own. Disclosure then adds a second helping, plus the static utilities, plus the chains that disappear from the call graph.

So the advice from the last paper stands and gets sharper. Keep the scoring function out of the agent's context and use it to watch. Ask for structure in words if you want structure, and budget for it. Then watch the test suite, because in this experiment it was the only instrument that noticed what the structural numbers were celebrating.


References and notes

  1. SlopCodeBench. No-context iterative extension, erosion and verbosity metrics, degradation slope. arXiv:2603.24755.
  2. SWE-CI. Continuous-integration gate, Normalized Change, EvoScore, Zero-Regression Rate. arXiv:2603.03823.
  3. F. P. Brooks Jr. No Silver Bullet: Essence and Accident in Software Engineering. Proc. IFIP Congress, 1986. Reprinted in IEEE Computer 20(4), 1987.
  4. L. Tesler. The law of conservation of complexity, articulated in the mid-1980s. Every process carries an irreducible complexity that design can move but not remove. Design decides only who bears it.
  5. C. A. E. Goodhart. Problems of Monetary Management: The U.K. Experience, 1975. The commonly quoted form is M. Strathern's restatement, 1997. When a measure becomes a target, it ceases to be a good measure.
  6. Change Impact in the Wild. An external test of the metric against defects in twenty human-written repositories. blog.officefloor.net.
  7. PetClinic-Evolve degradation study. Prior posts and harness. blog.officefloor.net.
  8. Cyclomatic complexity and per-function line ranges are computed with lizard. It is pinned and probed before each run. Clone detection uses jscpd at pinned thresholds. The smell pass is PMD against a ruleset that deliberately excludes every complexity rule.
  9. Telling the Agent the Cost Function. The disclosed-formula condition, run blind-202609010045, September 2026. blog.officefloor.net.

PREPRINT · OFFICEFLOOR · SEPTEMBER 2026

Wednesday, 16 September 2026

The decay you cannot see: what we measured and the one thing we will not claim

AI keeps the tests green while a codebase quietly rots. Here is how we measured that decay, what held up when we tried to break the metric, and why we refuse to sell it as a bug predictor.

The problem hiding behind green tests

AI is very good at complexity. That turns out to be the problem.

It keeps piling logic into code that already exists, because complexity does not slow it down the way it slows a person down. A god method grows another branch. A god class gains another method. The tests still pass. The feature still ships. Nobody notices.

Slowly the system decays to a point where no human can hold it in their head. Given enough of this, even the AI loses the thread, and every change starts breaking something else. By the time anyone feels the pain, the cheap fix is gone. What is left is an expensive, error prone refactor, or a rewrite.

The obvious question is whether you can see this coming. So we tried to measure it.

Score the change, not the code

Most quality metrics score a snapshot. They tell you how complex the code is right now. That is the wrong frame for decay, because decay is not a state, it is a motion. It is the act of adding to something already heavy.

So we measure the change, not the code. And we weight each change by how much existing tangled code it disturbs.

The core of the measure is the surrounding weight. That is the complexity already sitting in the class or module you are editing, measured on the state before your change. Adding a brand new file scores almost nothing, because nothing was there before. Growing an existing god class scores a lot. That asymmetry is the decay signal.

The full formula, and the three attempts it took to get there, are in an earlier post in this series. The short version is that a change costs more when it touches more files, disturbs more surrounding complexity, and adds more lines to already complex functions.

What the experiment showed

We ran two codebases through a long series of AI made changes. One was built additively, where you compose new pieces and wire them together. The other was built mutatively, where you edit a woven whole.

By the end of the run, a single change to the mutative codebase disturbed on the order of fifteen times more weighted structure than the same kind of change to the additive one. The decay was not hypothetical. It compounded, change after change, exactly as the thesis predicted.

Then we checked whether the number meant anything real, by correlating it against independent signals we captured during the run.

Signal Additive arm Mutative arm
Money cost of the change 0.69 0.54
Re-reading of existing code 0.60 0.49
Model time spent 0.68 0.53

The correlations held inside each arm independently, so this was not just an artifact of one codebase being harder overall. And when a change broke a rule in code it never touched, that change scored roughly seven to ten times higher than a change that broke nothing. The measure was tracking real cost, real comprehension load, and real blast damage.

This is the part where most tool posts would stop.

Then we tried to break our own metric

A measure that only ever flatters itself is worthless. So we took it to a wider, messier world. We scanned twenty open source repositories, across eight well sampled languages, and collected 349,165 per commit observations of the measure in the wild.

Then we asked the uncomfortable question. If this number is good, it should help predict where the bugs are. Does it?

It does not. Across that corpus, the measure does not beat plain file size as a defect predictor. File size is a famously strong and famously simple baseline for bug proneness, and our carefully weighted change measure did not clear it.

We could have quietly not run that test. We ran it, and we are telling you the result, because the result matters.

Why that is the right outcome

A negative result is only a failure if you were claiming the thing it disproves. We were not.

The measure was never built to predict defects. It was built to measure change effort and structural decay. Those are different quantities. File size predicts bugs well precisely because bigger files have more surface for anything to go wrong. That says little about how much a specific change grows the parts of your codebase that are becoming unmaintainable.

So we do not sell a bug predictor. We would be competing with file size and losing, and we would be lying. We measure the thing the experiment actually validated: how much effort a change costs, and how much it adds to the decay. That is a different and, we would argue, more actionable thing. It tells you where change is getting expensive, while it is still cheap to fix.

What it is actually good for

Three things.

It gates decay. When a change piles too much onto already heavy code, it can warn, or block the merge, and ask you to simplify or refactor first.

It routes attention. Every score comes with a ranked list of the files where the decay is concentrating, so a class quietly growing into a god class surfaces as a refactor candidate before it blocks anything.

And it grades on a curve, because a raw threshold cannot work. Change impact varies by around 300 times across projects and 60 times across languages. So instead of guessing a number, it grades a change by its percentile against a real distribution, blending a seed corpus with your own repository history.

The tool

All of this ships as ImpactGate. It runs as a command line tool, a git pre-commit hook, and as a check on GitHub, GitLab, and Jenkins. It treats an AI written change like any other change. It scores the blast radius, and the big ones stop for a human to look at while the fix is still small.

You can try it out with the following on your current code change:

docker run --rm -v "$PWD:/repo" ghcr.io/officefloor/impact-gate score

Honesty is the point

We think the interesting story here is not that we built a metric that tells a nice tale. It is easy to invent one of those. The interesting story is that we took our own metric to a place where it could fail, watched it fail at a job it was never meant to do, and kept the job it is good at.

If you are going to let AI write a lot of your code, you want tools that are clear about exactly what they measure, and honest about what they do not. This is a tool to highlight decay early, before it becomes expensive to fix.

Sunday, 13 September 2026

The Number the AI Never Saw: We Spent the Change Impact With a Tool

This is a series on how software architecture shapes the way AI-written code decays. Last time we did something deliberately bad. We handed the AI the exact cost formula we were judging it by, and told it to keep the number low. It did. The number fell about forty-five fold. The code got worse. It scattered logic into tiny classes and copied code instead of reusing it. It optimised the number, not the design.

The obvious lesson was do not show the AI the number. So this run keeps the same number. But the AI never sees it. A tool sees it instead.

We gave the number to a tool, not the AI

The setup is the same as always. Two codebases that do the same REST service. Spring puts each new rule into one growing controller. OfficeFloor spreads each rule across many small wired functions. Ten independent runs on each side. Sixty change requests per run, one after another. Add a validation rule. Change how a field is stored. And so on for sixty steps.

The AI gets the plain request and nothing else. No formula. No cost. No hint that anything is being measured. It just implements the change.

Behind it sits the gate. On every change the gate scores the structural impact with the same formula as before. Here is that formula in plain terms.

cost = (complexity of the rest of the class) x (complexity of this method) x (lines changed)

The big term is the first one. Editing a method in a large class makes you carry the whole class in your head. So a change that piles more into an already heavy class scores high.

What the tool does when it fires

When a change scores over the line, the tool does not lecture the AI. It quietly runs one refactor first, on the clean code, before the change goes in. It tells the refactor step only the names of the classes that got too heavy. It does not say the cost. It does not say the formula. It does not say the words split into small classes. Those exact words are what caused the mess last time. It just asks for the code to be prepared so the next change lands in a way a maintainer would find natural.

Then the change is attempted again, on the cleaner code. The AI writing the feature never sees a score. There is nothing for it to optimise, so there is nothing for it to game. The number only ever points the tool at where to cut.

It worked on the change in front of it

The run finished clean. All twenty runs reached all sixty steps. Nothing broke the build. Nothing stopped early.

The tool fired on about a hundred and seventy Spring changes and about two dozen OfficeFloor changes. Spring trips far more often, which fits the whole series. Spring concentrates. OfficeFloor does not.

And when the tool fired, the change landed cleaner. The typical flagged Spring change dropped from an impact of about 11,800 to about 8,100. The typical flagged OfficeFloor change dropped from about 10,700 to about 5,100. So the mechanism is honest. The tool catches the change that would concentrate complexity, and moves that complexity somewhere less painful before the change goes in.

Then we looked at the standing code

A cheaper change on the day is not the point. The point is the shape of the code at the end. So we measured the heaviest class in the app, on the final code, the same way for every run. The AI could never aim at this number, because it never saw it.

Spring's biggest class at the end Weighted complexity
Plain spec, no gate135
Told the formula22
Hidden gate with refactor96

The hidden gate brings the god class down from 135 to 96. A real drop. Not a cure. The formula run shows 22, which looks far better. Hold that thought.

The 22 was a lie

The formula run scored beautifully because the AI gamed the score. So we counted what it actually built.

Spring, per run Plain spec Told formula Hidden gate
New classes made93113
Tiny static helpers2215
Duplicated lines2,2302,5102,230

The formula run tripled the class count. It made ten times as many tiny static helpers. A static helper in a small class costs almost nothing in the formula, so it is a cheap place to dump logic. It also duplicated the most code, because reuse means editing a class the formula was punishing. The god class number went down because the logic moved out of the controller and into a crowd of little files. A static helper costs you things the number does not see. No dependency injection. Harder to mock in a test. No place in a transaction.

The hidden gate did none of that. Its class count, its helper count, and its duplication all sit right next to the plain baseline. It lowered the real concentration by moving real structure. It did not dodge the number by hiding the complexity somewhere the number cannot look.

The total complexity never actually dropped

Here is the number that ties it together. We added up all the complexity that ended up in the code, wherever it lived.

Spring, total complexity added Amount
Plain spec, no gate282
Told the formula268
Hidden gate with refactor278

The total is about the same in all three. The work to build the feature does not shrink because you measured it. All any of these runs can change is where the complexity sits. The formula run shoved it into a pile of new files and called the job done. The hidden gate spread it a little more sensibly across the real code. Neither made it go away.

This is a very old idea

Fred Brooks split complexity into essential and accidental. The essential part is the difficulty of the problem. You cannot delete it. Larry Tesler said every system has an amount of complexity that cannot be removed. The only question is who carries it. That is exactly what the total above shows. And Goodhart said a measure that becomes a target stops being a good measure. That is exactly what the formula run showed. The hidden gate is the way to use a measure without making it a target.

Spring still grew a big class

Be clear about the limit. The hidden gate helps. It does not fix. Spring still ended with a class at 96 weighted complexity. That is still a big class. The tool lowers the impact of the change in front of it, but the complexity it pushes off today lands on a later change. A few Spring runs still ran away to a full god class anyway. And the gate did not make the code safer. It delivered a touch less and broke a touch more than the plain baseline. The win is structural. It is not free.

What to take from this

  • Do not hand an AI the metric you are judging it by. It will optimise the metric and not the code.
  • A metric is still useful. Let a tool read it and spend it. Keep it away from the thing writing the code.
  • When you do spend it, aim a refactor at the sore spot. Name the heavy class. Do not prescribe the fix.
  • Do not expect the complexity to vanish. Expect to choose where it lives.

What we are not claiming

This is one model, one kind of task, two codebases. The gate here records and continues. It does not throw changes away. That is a different test. A fourth run is going now. It gives the AI plain English advice to write good structure, with no number at all. That will tell us whether a tool beats simply asking nicely. Until then, the honest summary is small. Keep the number for the tool. Keep it away from the AI. Point it at the sore spot and let it cut.

Spending the Cost Function Without Telling the Agent

Preprint / cs.SE / Empirical Software Engineering

Spending the Cost Function Without Telling the Agent

Two architectures. Twenty chains. Sixty accumulating rules each. A structural metric read by a tool instead of handed to the agent.

OfficeFloor · independent research · blog.officefloor.net

September 2026 · correspondence: daniel@officefloor.net


Abstract

Change impact scores a code change by the complexity it disturbs. A prior study in this series disclosed that cost function to a coding agent as its objective. The agent optimised the number and not the code. That was Goodhart's law, and it held cleanly. This paper keeps the same cost function but never shows it to the agent. The agent receives a plain specification. A tool scores each change in the background. When a change concentrates too much complexity, the tool runs one refactor first, on the clean code, and names only the heavy classes.

Across two architectures and twenty independent chains of sixty accumulating changes, the hidden gate lowers the rate of structural concentration by a real amount. Spring's per-checkpoint impact slope falls from 435 to 127, a cut of about seventy percent, with non-overlapping intervals. It does so without the class proliferation and duplication the disclosed formula produced. The disclosed run reported an impact slope of 9.7 for Spring, far lower, yet it tripled the class count, multiplied static utilities nine fold, and broke the most behaviour.

The total complexity that lands in the codebase is about the same in all three conditions. Only its distribution moves. Disclosure moved it to where the measure was blind. The hidden gate moved it in the code. The reduction is partial, the functional cost is small, and Tesler's conservation of complexity still holds. The metric works as an instrument a tool spends. It does not work as a target handed to the writer.

Keywords: Goodhart's law · conservation of complexity · AI-assisted development · software architecture · code degradation · refactoring · cyclomatic complexity

1Introduction

A prior post handed the agent the exact cost function and told it to minimise it. The number fell by about forty-five fold. The code did not improve. The agent scattered logic into tiny classes where the surrounding cost is near zero. It copied code rather than edit a class the formula already penalised. The concentration the metric was meant to prevent simply moved to where the metric was not looking.

This raises a narrower question. The failure above was disclosure, not the metric. So can the same cost function reduce concentration if the agent never sees it, and a tool spends it instead. The intervention here is a gate that reads the score and acts on it, with a wall between the tool and the agent. The agent writes the code. The tool watches and, when needed, prepares the ground.

2Background and related work

Brooks separated essential complexity, the difficulty of the problem, from accidental complexity, the part we add [1]. Essential complexity cannot be deleted by better instructions. Tesler's law of conservation of complexity says every system carries an irreducible amount [2]. The only open question is who holds it. Goodhart's law says a measure that becomes a target stops being a good measure [3]. The disclosed run was a direct demonstration of all three at once.

The structural measures are standard. Cyclomatic complexity follows McCabe [4]. Weighted methods per class follows Chidamber and Kemerer [5]. The cost function combines them to price the context a change must disturb [6].

3The metric and the intervention

For every function a change touches, the cost is:

cost = max(WMC_other, 1) × CC × max(1, changed_lines)

CC is that function's cyclomatic complexity. WMC_other is the summed complexity of the other methods in its class, the context a maintainer must hold to edit it safely. changed_lines is how many of its lines the change adds or edits. The costs are summed over every changed function and multiplied by the number of files touched. The dominant term is WMC_other. A method inside a large class pays for the whole class.

The gate scores each change against a fixed percentile of a reference distribution. A change over the line is flagged. On a flag the tool does not lecture the agent. It runs one refactor step on the clean pre-change code. The refactor prompt names the heavy classes and asks for a natural restructuring. It never states the cost, the formula, or the phrase "split into small classes". Those phrases drove the dispersal and duplication in the disclosed run.

A change is about to be made to this codebase. When that change was
implemented directly, these classes ended up carrying more complexity
than they can hold comfortably:

{drivers}

Before the change is made, prepare the existing code so the change can
land as code a maintainer of this codebase would recognise as natural,
consistent with how this codebase already structures similar work.

A deterministic quality gate checks the refactor's own added lines for duplication before the change is re-attempted. The gate is advisory. It records the outcome and continues. A flag never aborts a chain. The number only ever points the tool at where to cut.

4Study design

Two codebases implement the same REST service. Spring routes each new rule through one growing controller. OfficeFloor spreads each rule across many small wired functions. The same total complexity lands in both, but Spring concentrates it. This is the independent variable the whole series holds.

Each arm runs ten independent chains. Each chain applies sixty sequential change requests to the service. Every arm gets the same specifications and the same plain implement prompt. The agent is the same model throughout. The gate threshold is a shared cutoff, calibrated to OfficeFloor's own cohesion, so OfficeFloor trips rarely and Spring trips often. All structural measures are recomputed on the standing code at each checkpoint, the same way for every condition, on numbers the agent could never aim at. Slopes are ordinary least squares of the metric on checkpoint index. Confidence intervals are ninety-five percent, bootstrapped over chains.

Three conditions are compared. Control is the plain specification with no gate. Disclosed is the earlier run with the formula in the prompt and a hard gate. Refactor-only is the hidden gate described above. A fourth condition, plain English advice with no number, is still running and is not reported here.

5Results

The refactor-only run completed cleanly. All twenty chains reached the final checkpoint. No build broke. No chain stopped.

When the tool fires, it works on the change in front of it. The refactor triggered on 171 of 600 Spring checkpoints and 24 of 600 OfficeFloor checkpoints. On a flagged change the accepted impact is lower than the first attempt. The median flagged Spring change falls from about 11,800 to about 8,100. The median flagged OfficeFloor change falls from about 10,700 to about 5,100. So the per-change mechanism is honest. Spring trips the gate far more often, which is the concentrated versus cohesive split the series keeps finding.

The result that matters is the standing code. Table 1 gives the rate at which the scored impact of each new change grows over a chain.

Change-impact slope per checkpointControlDisclosedRefactor-only
Spring4359.7127
OfficeFloor765.728

Table 1. Ordinary-least-squares slope of change impact on checkpoint index. Lower is slower concentration. Spring refactor-only 127 [76, 179] versus control 435 [307, 587]: the intervals do not overlap.

Table 2 gives the standing god class. This is the weighted complexity of the class the endpoint routes through in both arms, measured on the final code, and its growth per checkpoint.

Handler class, SpringControlDisclosedRefactor-only
Final weighted complexity13522.596
Growth per checkpoint1.770.041.17

Table 2. The hidden gate dents the god class but does not cure it. Spring still ends with a single class at about 96 weighted complexity, below the control's 135 and far above the disclosed run's 22.

The disclosed number looks like a win until you look at where the code went. Table 3 is the gaming signature on Spring, the arm with a monolith to dodge.

Spring, per chain unless notedControlDisclosedRefactor-only
New classes created931.513.3
Of those, static utilities2.320.65.2
Duplicated lines at tip2,2302,5102,230
Cumulative CC, in new files4621899
Cumulative CC, in existing files23650178
Cumulative CC, total282268278

Table 3. Read the last row first. The total complexity is about the same in all three. Only its distribution moves. The disclosed run inverted where the complexity lives and raised duplication. The refactor-only run sits at or near the control on every row.

Table 4 gives the cost. EvoScore is a functional delivery score over the chain. True regressions count previously passing behaviour the agent later broke, excluding intended changes.

SpringControlDisclosedRefactor-only
EvoScore0.7870.4400.712
True regressions375641

Table 4. The disclosed run was also the worst at the task. The refactor-only run stays close to the control. The structural gain does not come with a safety gain.

6Discussion

The refactor-only gate gives a real reduction in concentration, and an honest one. The standing evidence shows structure moved, not a number gamed. But the reduction is partial. The per-change cut is reliable, yet the standing curve bends only part way. The complexity the tool pushes off today's change lands on a later one. That is conservation again. A gate on each change slows the concentration on the arm that concentrates. It has not stopped it.

The contrast with disclosure is the core finding. Disclosure produced a far lower score and worse code. The hidden gate produced a higher score and better-shaped code. The same cost function, read by a tool rather than chased by the writer, changes the structure instead of the number.

7Threats to validity

One model, one task family, one pair of architectures. The gate here is advisory, so it records and continues rather than discarding a change. A hard gate is a separate condition. The threshold is calibrated to OfficeFloor's cohesion, which fixes how often each arm trips. Ten chains per arm leave real variance, and Spring's variance is wide. Some chains still run away to a full god class. The smell half of the duplication and pattern detector contributed a negligible number of lines, so the duplication figure is effectively a clone measure. EvoScore and true regressions are measured independently of the structural score, which is why they can disagree with it.

8Reproducibility and data availability

Availability

Each chain is a git branch of sequential checkpoint commits. The analysis recomputes every structural measure from the commits, so no derived value is read back. Each run pins its own configuration snapshot, so the metrics match how that run was scored. Confidence intervals are bootstrapped over chains.

9Conclusion

Brooks's essential complexity did not shrink under a plain spec, a hidden gate, or a disclosed formula. Tesler's conservation held in every condition. The total complexity landed in roughly the same amount each time. When the measure was disclosed, the work moved to wherever the measure was not looking, and the code got worse. When the measure was kept back and spent by a tool, the work moved in the code, and the concentration fell part way with no gaming. Use the metric to watch and to aim a refactor. Do not hand it to the writer as the goal.


  1. F. P. Brooks. No Silver Bullet: Essence and Accidents of Software Engineering. Computer, 1987.
  2. L. Tesler. The Law of Conservation of Complexity. Mid-1980s.
  3. C. A. E. Goodhart. Problems of Monetary Management: The UK Experience. 1975.
  4. T. J. McCabe. A Complexity Measure. IEEE Transactions on Software Engineering, 1976.
  5. S. R. Chidamber, C. F. Kemerer. A Metrics Suite for Object Oriented Design. IEEE Transactions on Software Engineering, 1994.
  6. D. Sagenschneider. Telling the Agent the Cost Function. blog.officefloor.net, September 2026.

Slopes are OLS on checkpoint index · intervals 95% bootstrapped over chains · ten chains per arm · sixty checkpoints per chain

Monday, 7 September 2026

The refactor run halfway

The last post set up a ladder. Four runs, each adding exactly one thing. Plain spec, the baseline erosion. Plain English, good advice with no number. The refactor run, a tool that catches the sore spot and nudges it, still no number shown. The formula, the full measurable target, already run and already gamed.

This is an update on the third rung. The refactor run is about halfway through, and it is behaving. So here is what I can already see, and what I still cannot.

What the refactor run is

It is the tool lever on its own. The agent gets the plain spec and nothing else. No formula, no cost, no mention of the experiment. It just implements the change.

Behind it sits ImpactGate. On every change it scores the structural impact with the same cost formula from the earlier posts. When a change scores above the line, the tool does not lecture the agent. It runs a refactor step on the clean code, names the heavy classes, and breaks them into smaller cohesive ones. Then the change is attempted again on the cleaner code.

The key difference from the formula run is the wall between the tool and the agent. The agent never sees the number. There is nothing for it to optimise, so there is nothing for it to game. The number only ever points the tool at where to cut.

How far it has run

Two codebases, Spring and OfficeFloor, the same two as always. Ten independent chains each, sixty checkpoints a chain. Six chains are done on each side, and the seventh is running. So a little over half the run is in.

It is clean so far. Every finished chain ran all sixty checkpoints. No build ever broke. No chain stopped early. The gate is advisory here, so it records and continues rather than aborting, which is why a flag never derails a chain.

The one thing I can already see

When the tool fires, it works.

The refactor has triggered on about a hundred and twenty checkpoints so far. Every one of them was a change the formula scored near the top of the scale before the cut. After the refactor, the same change lands on cleaner code and its impact score falls by about two thirds. The average impact on those changes drops from roughly nineteen thousand to roughly six thousand. It falls in about nine of every ten cases. Spring trips the gate far more often than OfficeFloor, which is exactly the concentrated-versus-cohesive split the whole series keeps finding.

So the mechanism is honest. The tool catches the change that would concentrate complexity, and the refactor moves that complexity somewhere less painful before the change lands. No number was ever shown to the thing writing the code.

The curve it has to bend

A cheaper change on the day is not the point of the experiment. The point is the shape of the standing code at the end. So the number that matters is the heaviest class in the app, measured by its weighted complexity, and how fast that grows checkpoint after checkpoint. That is a god class forming in slow motion. It is measured on the final code, the same way for every rung, on a number the agent could never aim at.

On the plain spec baseline the biggest class keeps growing. Spring's heaviest class gains about one and eight tenths of weighted complexity every checkpoint. OfficeFloor's gains about seven tenths. Spring concentrates roughly two and a half times faster. That gap is the erosion the tool is meant to fight.

An early read of the slope

I did not want to quote a number off half a run. But half a run with nothing in it says very little, so here is the honest interim, with the caveats loud. Six of the ten chains are in on each side. Measured the same way as the baseline, the growth per checkpoint so far is:

  • OfficeFloor. Baseline about seven tenths. Refactor run about six and a half tenths. Essentially unchanged.
  • Spring. Baseline about one and eight tenths. Refactor run about one and a half. Lower, but the chains are all over the place.

OfficeFloor is the easy read. It was already cohesive, so the tool rarely fires and there is little to bend. Its slope barely moves, which is what you would expect when the problem was never there.

Spring is the interesting one, and not in the clean way I hoped. The average slope drops by about a fifth. But the spread is wide. Four of the six Spring chains stayed reasonably flat. Two of them still ran away, ending with a single class carrying about a hundred and sixty weighted complexity, a full god class, gate and refactor notwithstanding.

So the per-change cut is real, and it is not reliably reaching the end state. The tool lowers the impact of the change in front of it. The standing god class on the arm that has the problem still grows at roughly four fifths of the baseline rate, and unevenly. The complexity the tool pushed off today's change is landing on some later one.

Treat those numbers as a direction, not a verdict. They are six chains, not ten. The final figure uses a more careful slope with confidence bands, and Spring's variance is exactly the kind that moves once the last chains land. The sign looks right. The size is not settled.

Where this leaves the ladder

The formula run proved a measurable target gets gamed. This run is the opposite bet. Keep the number away from the agent, and let a tool spend it instead.

Halfway in, the tool does its job on every change it touches, and the standing curve bends only part way. That is the Tesler shape again. The complexity did not leave. It moved. A gate on each change slows the concentration on the arm that concentrates, but it has not stopped it, and twice in six tries it barely dented it.

A few more days and both sides finish. Then the real measurement, on the standing code.

Sunday, 6 September 2026

We told the AI about the Change Impact

A series on how software architecture shapes AI-driven code degradation. This is the plain version of a result. The paper has the numbers and the caveats.

We have been running an experiment. An AI agent gets sixty change requests, one after another, all landing on the same REST endpoint. Add a validation rule. Change how phone numbers are stored. Add a duplicate check. Sixty of them, in order, with the full test suite run after each one.

We do this twice. Once on a normal Spring codebase. Once on the same application built with OfficeFloor. Then we look at what sixty changes did to each one.

The Spring result has been the same every time. The create-owner handler method grows. Rule after rule lands in it. By the end it is the biggest thing in the codebase and every new rule means opening the same enormous method again.

And there has always been an obvious objection to that. Nobody told the AI not to do it.

So we told it

We have a metric called change impact. It scores a change by how much complexity it disturbs, rather than by how many lines it edits. Touching a small method in a small class is cheap. Adding ten lines to a huge method in a huge class is expensive.

Roughly, for every method you change:

cost = (how heavy the class already is)
     x (how complicated the method is)
     x (how many of its lines you changed)

Add that up for every method you touched. Multiply by the number of files. Lower is better.

This time we put that formula in the prompt. Every single change request. We told the AI exactly how it was being scored, and we told it which part of the formula mattered most.

It worked immediately

We had also built a safety net. If a change scored badly, we would throw it away and make the AI refactor first. We expected that to be doing the work.

It fired three times. Out of twelve hundred changes.

The prompt alone was enough. Told what the score was, the AI wrote code that scored well on it, first try, nearly every time.

And the god method never appeared. Here is how much the controller file grew over sixty rules:

Spring controller, lines added over 60 rulesTen runs
Normal prompt604 to 962
Told the formula2 to 44

That is the objection landing. Prompt better and the problem goes away. If we stopped here, this post would say architecture does not matter much and good instructions do.

Then we looked at where the code went

We have a second check that does not care about any of our metrics. It takes the finished codebase, finds every line the run changed, and works out which method that line ended up in. Then it adds up how complicated all those methods are. It is deliberately dumb. It just asks how much logic now exists and where it lives.

Spring, after 60 rulesNormal promptTold the formula
Total logic written282268
...sitting in brand new files46218
...sitting in files that already existed23650
Number of files involved1941

Look at the first row. The amount of logic is the same. 282 before, 268 after. Nothing got simpler.

Now look at the next two rows. It all moved. Out of the files that existed, into files the AI created. Twice as many files.

This is not the AI being clever or sneaky. Look at the formula again. The first term is how heavy the class already is. A brand new file has nothing in it. So that term is as small as it can possibly get. If you want a low score, the cheapest thing you can do is put your code somewhere nothing else lives.

So it did. Sixty times.

What thirty-one new classes look like

Per run, the normal prompt created about nine new classes. Told the formula, it created about thirty-one. Roughly twenty of those were static utility classes. Classes with one static method, holding one rule, and nothing else.

On the score, that is perfect. A static method in an otherwise empty class has almost no surrounding complexity to pay for.

In a real Spring codebase, it costs you things a junior engineer runs into fast. A static method is not a bean. You cannot inject anything into it. You cannot swap it out in a test. Spring cannot wrap it in a transaction or a proxy. You have made the metric happy and given up most of what the framework is for.

The duplication went up too, and for a reason worth understanding. Reusing an existing helper means adding a line to a method in a class that already has weight. The formula charges you for that. Writing your own copy in a fresh file is free. So the AI wrote its own copy. Duplicated lines went from about 2230 to about 2510, in a codebase that had got smaller overall.

One run vanished completely

We ran ten independent Spring chains. Most of them dispersed into helper classes as described. Two did something else.

In one of them, the create-owner handler at the end of sixty rules is the method we started with. Map the request. Save the owner. Set the location header. Return 201. That is it.

Every one of the sixty rules is a @RestControllerAdvice class. Eighteen of them. They run before the handler is ever called, because Spring invokes them, not the handler.

Our comprehension metric follows method calls from the endpoint. It is a good metric. It was added because a reader correctly pointed out that a pipeline can hide work in its later stages, and following the calls fixes that.

It cannot follow something nothing calls. An interceptor is invoked by the framework. So for that run, our metric reports the create path as having a complexity of 3, for a codebase implementing sixty business rules.

That is not a clean codebase. If you are asked to change the phone number rule, you still have to find it first, and finding it just got much harder. The number got quieter. The code did not get simpler.

The part that actually matters

All of the above is arguing about metrics. This part is not.

SpringNormal promptTold the formula
Implemented the rule it was asked forevery timeevery time
Whole test suite still green79% of changes44% of changes
First rule permanently broken atrule 47rule 24

The AI still did the job in front of it. Every time. What it stopped doing was keeping the previous fifty-nine rules working.

Something broke, earlier and more often, and none of the structural metrics showed it. The tests showed it.

We had a theory. Rules scattered across separate interceptors still have to run in some order, that order is no longer written down anywhere, and an AI adding rule 40 cannot see the ordering it is joining. It is a nice theory. The finished data does not support it. The runs that leaned hardest on interceptors actually broke slightly less. So we do not know why yet, and we are saying so rather than keeping a tidy explanation that the numbers disagree with.

This is a very old idea

Two of them, actually.

Fred Brooks, in 1986, split the difficulty of software into two parts. Essential complexity is the problem itself. Sixty business rules are sixty business rules. Accidental complexity is the mess we add on top through how we choose to build it. His argument was that no tool removes the essential part.

That is exactly the first row of our second table. 282 before, 268 after. The rules are the rules. No prompt made them cheaper.

Larry Tesler put it a different way in the 1980s, usually called the law of conservation of complexity. Complexity does not disappear. It moves. Design decides who has to deal with it, not whether anybody does.

That is the rest of the table. The complexity moved out of the handler and into forty-one files, and in two runs it moved somewhere our tooling could not follow at all.

And the reason it moved is Goodhart's law, in its usual form: when a measure becomes a target, it stops being a good measure. We knew that. We still did not expect it to happen this completely, in one run, from one paragraph of prompt, with no gate ever firing.

What to take from this

If you are early in your career and working with AI tools, this is the practical version.

A green metric is not the same as good code. When a number improves a lot and quickly, ask what moved. Not what got deleted. Things rarely get deleted.

Be suspicious of a class that exists to hold one rule and nothing else. It is often a real improvement. It is also the cheapest way to make almost any code metric look better, so it is worth checking which one you are looking at.

Static helpers look free and are not. You give up injection, mocking, transactions and proxying. If someone, human or AI, is producing a lot of them quickly, that is worth a conversation.

If you cannot find where a rule runs, the codebase got harder, whatever the dashboard says. Being able to open one method and read what happens is worth a lot.

And keep your tests. In this experiment every structural metric we own said the code got better. The acceptance suite was the only thing that noticed rules quietly breaking from change 24 onwards. That is not a small detail. It is the whole reason we caught this.

What we are not claiming

We told the AI to optimise a formula and things got worse. We have not yet shown that a normal request for good structure would do the same. There is a difference between "here is the arithmetic, minimise it" and "please keep this code well organised", and we are running that second version now.

Until that finishes, the honest claim is narrow. Do not hand an AI the metric you are judging it by. Use the metric to watch. Do not use it as the goal. Those are two different jobs and it can only do one of them.