Sunday, 6 September 2026

Telling the Agent the Cost Function

Preprint / cs.SE / Empirical Software Engineering

Telling the Agent the Cost Function

Two architectures. Twenty chains. Sixty accumulating rules each. What happens to a codebase when its structural metric becomes the objective.

OfficeFloor · independent research · blog.officefloor.net

September 2026 · correspondence: daniel@officefloor.net


Abstract

Change impact scores a code change by the complexity it disturbs [6]. It was defined inside a controlled degradation study. There it tracked how a codebase erodes as an AI agent lands accumulating changes on one endpoint. This paper reports what happens when that score is handed to the agent as its objective. Ten independent chains per architecture. Sixty accumulating rules per chain. The same endpoint and the same acceptance suite as an untouched control.

Disclosure works on the disclosed measure. The per-checkpoint impact slope falls 45-fold for Spring. It falls 13-fold for OfficeFloor. Spring's growing handler does not appear. That handler is the central finding of the prior work in this series. Its controller file grows by 604 to 962 lines across the ten control chains. It grows by 2 to 44 lines across the ten disclosed chains. Two of the three concentration statistics stop separating the architectures at all.

The complexity did not leave. We attributed every changed line at the chain tip to the function that now contains it. The total cyclomatic complexity Spring's run touched is 282 before and 268 after. That is unchanged inside its spread. The share living in files the run created rises from 46 to 218. The file count doubles. The rules were relocated. They were not removed. This is Brooks's essential complexity [3], conserved as Tesler described [4]. What rose is the accidental part. Created classes per Spring chain go from 9.0 to 31.5. Two-thirds of them are static utilities. Duplicated lines rise on a codebase that shrank. In two of ten chains the rules moved into framework-dispatched interceptors. Those rules left the call-graph measurement entirely. One chain reports a create-path complexity of 3 with all sixty rules implemented.

Delivered correctness fell. Every checkpoint's own new rule still landed in every chain. func is 1.000 throughout. Retention of previously passing rules dropped from 0.787 to 0.440 for Spring. It dropped from 0.732 to 0.577 for OfficeFloor. The first standing failure arrives at rule 24 instead of the high forties. We report the mechanism as unexplained. The ordering hypothesis we proposed mid-run is not supported by the completed data. The conclusion is narrow and firm. Change impact is usable as evidence about code that is not targeting it. It is not usable as an optimisation target.

Keywords: change impact · Goodhart's law · conservation of complexity · AI-assisted development · software architecture · code degradation · metric gaming · cyclomatic complexity

1Introduction

Every result in this series has had the same rebuttal waiting for it. Spring's request handler accumulates rule after rule. It ends as the largest thing in the codebase. OfficeFloor's wired pipeline stays flat. Of course it does. Nobody told the agent not to let it happen.

So we told it. This paper reports a full sweep in which the agent is handed the exact arithmetic its work will be scored by. It is handed it at every one of sixty checkpoints. Not advice about clean code. The cost function itself, with its dominant term named.

The question is not whether an agent can optimise a disclosed objective. It can. The first result below is how completely. The real question is what optimising it does to the code underneath. We measure that with instruments the agent was never told about. An unscoped audit of where complexity physically ended up. A classification of what kind of class now holds each rule. A duplication detector. The acceptance suite.

2Background and related work

That an optimised measure stops measuring is old. Goodhart observed it for monetary policy. Strathern's restatement is the one usually quoted [5]. When a measure becomes a target, it ceases to be a good measure. The software-metrics literature has its own long record of this. Lines-of-code targets are the familiar case. What is new here is the speed, and the operator. An AI agent given a formula optimises it immediately. It does so at every checkpoint. It does not tire and it does not negotiate. So the failure mode becomes observable inside a single controlled run. It no longer needs quarters of organisational drift to show up.

Two older results frame what we found underneath. Brooks divided the difficulty of software in two [3]. Essence is the complexity inherent in the problem being solved. Accident is the complexity introduced by how we happen to build it. He argued that no tooling improvement removes the essential part. Tesler's law of conservation of complexity makes a related point about placement rather than tooling [4]. A given task carries an irreducible amount of complexity. Design decides who bears it. Design does not decide whether it exists. Tesler's framing is usually applied across the user and developer boundary. Our result is an instance of it inside a single codebase. Sixty business rules are essential complexity. The formula moved them. It did not remove them. The accidental part grew.

The experiment's measures come from two sources. SlopCodeBench supplies erosion, verbosity and degradation slope [1]. SWE-CI supplies normalized change, EvoScore and zero-regression rate [2]. The change-impact score itself is defined and validated elsewhere [6]. That validation tests it against defects in human-written code.

3The metric and the intervention

Change impact charges a change by the complexity it disturbs. It does not charge by the lines it edits. For each function the change adds lines to:

cost(function) = max(WMC_other, 1) × CC × max(1, Δlines)

CC is the changed function's own cyclomatic complexity. Δlines is how many of its lines were added or edited. WMC_other is the summed complexity of the other methods in its class. It is measured on the state before the change. It stands for the surrounding context a maintainer must hold to edit the function safely. The per-function costs are summed. The total is then multiplied by the number of files touched. The full definition and rationale are in [6].

The intervention is simple. That text goes into the implement prompt at every checkpoint. The dominant term is named outright:

Write it to keep structural impact low. Structural impact is a
measurable cost, not a vague goal. For every function you change,
the cost is:
    cost = max(WMC_other, 1) * CC * max(1, changed_lines)
...
So the biggest lever is WMC_other: a method inside a large class
pays for the whole class. Do not grow one method or one class.

The control prompt is the instruction used unchanged throughout this series. Implement the specification. Make the tests pass. Nothing else differs between the two conditions.

4Study design

Research questions

  • RQ1. Does disclosing the cost function reduce the measured change impact, and the concentration statistics built on it?
  • RQ2. Does that reduction correspond to less complexity existing in the codebase?
  • RQ3. If not, where does the complexity go? Can the measurement still see it?
  • RQ4. What does the intervention cost in delivered correctness?

Arms, chains and checkpoints

The coding agent is held fixed. Architecture is the independent variable. One arm is a Spring @RestController codebase. The other is an OfficeFloor codebase of YAML-composed functions. Both implement the same PetClinic REST application. Sixty change specifications land in sequence on the single endpoint POST /api/owners. Most are additive. Every fourth from the eighth is mutative. A mutative checkpoint revises prior rules and ships updated copies of the affected tests. There are ten independent chains per architecture per condition. Each condition is therefore 1200 agent turns. The disclosed run is blind-202609010045. The control is blind-202608100006. Both use the same specifications and the same acceptance suite. Neither was modified between the two runs.

Blind protocol

The agent sees the current specification. It sees the accumulated tests up to and including the current checkpoint. It never sees future ones. Each turn runs in a history-less sandbox rebuilt from the worktree. So the agent cannot infer which checkpoint it is on from git history. Each turn also gets a fresh configuration and memory. Acceptance tests are injected per checkpoint rather than pre-committed. After the turn the full accumulated suite runs. Regressions are computed against the set of tests passing before the checkpoint.

The gate did not act

The disclosed run also carried an active gate. It could discard and re-attempt any change scoring above the 95th percentile of a reference distribution. It fired on 3 of 1200 checkpoints. One was OfficeFloor and two were Spring. It stopped no chain. All twenty chains reached checkpoint sixty. The mean accepted grade was the 31st percentile for OfficeFloor and the 38th for Spring. The prompt alone moved the agent so far below the threshold that the control loop had nothing to do. Everything reported below is therefore a prompt effect. We treat the condition as prompt-only.

Measures

Degradation slope is the OLS slope of a metric on checkpoint number. Intervals are 95% and come from a bootstrap clustered on chains. The between-architecture test is the difference of those slopes. Alongside the scoped metrics we compute an unscoped cumulative audit. That is one diff per chain, from the branch base to its final commit, over every changed file. Each changed line is attributed to the function containing it at the tip. We then sum the complexity of the distinct functions touched. It is scope-free. That makes it the check on every scoped number. Created classes are classified from the parsed function list. Duplication is measured by clone detection.

5Results

RQ1: disclosure works, on the disclosed measure

slope per checkpointSpring controlSpring disclosedOfficeFloor controlOfficeFloor disclosed
impact_composite435.19.7176.45.74
impact_mutation241.93.1717.62.12
impact_godclass193.26.5458.83.62
wmc_handler1.7690.038-0.0000.055
entry_cc0.1220.035-0.0000.036
erosion_handler0.00326000
node_cc_median3.0281.1010.0930.066

Table 1. Degradation slopes, control against disclosed, both architectures. Spring's change-impact slope falls 45-fold. Its mutation term falls 76-fold. wmc_handler is the weight of the class the endpoint routes through. entry_cc is the entry handler's own complexity. node_cc_median is the complexity reachable from one handling node. All are per-checkpoint OLS slopes over ten chains.

The difference of slopes between the two architectures is the actual test. It moves further. Under the control prompt, Spring minus OfficeFloor is +1.769 [1.597, 1.929] for wmc_handler. It is +0.00326 [0.00160, 0.00489] for erosion_handler. Both exclude zero. Under disclosure they become −0.018 [−0.054, 0.018] and exactly zero. The two architectures stop being distinguishable on the god-class statistics. The change-impact difference falls from +358.7 [234.2, 507.2] to +3.97 [0.78, 7.79]. That is a ninety-fold compression. It still excludes zero, but only just.

The plainest number is not a slope. Spring's controller file grows by 604 to 962 lines from base to tip across the ten control chains. It grows by 2 to 44 lines across the ten disclosed chains. The god method that this series was built on does not appear. Taken alone, that is the rebuttal landing. Prompt better, and the architectural difference goes away.

RQ2: the complexity is conserved

base to tip, per chainSpring controlSpring disclosedOF controlOF disclosed
CC sum over touched functions282.3 ± 22.2267.9 ± 41.5341.9 ± 13.2269.9 ± 22.8
  in files the run created46.3218.2286.2232.0
  in pre-existing files236.049.755.737.9
distinct functions touched110.3133.5145.7123.7
files parsed18.7 ± 5.441.4 ± 3.366.8 ± 4.866.6 ± 4.7

Table 2. The unscoped cumulative audit. Every changed line at the chain tip is attributed to the function that contains it. A function counts once, however many checkpoints edited it. Spring's total is unchanged inside its spread. Its distribution inverts and its file count doubles. This is the check no prompt-side scoping can evade.

Spring's total touched complexity is 282 before and 268 after. That is flat. It sits well inside the chain-to-chain spread. What changed is where it sits. Complexity in pre-existing files falls from 236 to 50. Complexity in newly created files rises from 46 to 218. The number of files involved doubles. WMC_other is the formula's dominant lever. A function in a brand-new file has no prior neighbours. Its WMC_other floors at one. So the cheapest way to satisfy the formula is to put the rule somewhere nothing else lives.

This makes Brooks's distinction measurable. The sixty rules are essential complexity. The problem requires them. No prompt made them cheaper. What the formula could change was their placement. That is Tesler's point about conservation, applied inside a codebase rather than across the user and developer line. One caution is worth stating. OfficeFloor's total did fall, from 342 to 270. That is a genuine reduction of about a fifth. We do not attribute it to relocation. It is Spring's total, the architecture under pressure, that is conserved.

RQ3: where the rules went, and what the measure could see

Spring, per chaincontroldisclosedrange, disclosed
classes created9.031.526–37
  static utility2.320.612–31
  injected bean0.07.30–25
  exception6.41.60–10
duplicated lines, final22302510 
production Java lines, final27002550 

Table 3. What replaced the god method. Two-thirds of the classes Spring now creates are static utilities. They score near-zero WMC_other. They also give up dependency injection, test seams, proxying and transaction participation. Duplicated lines rise on a codebase that shrank. Reuse means editing an existing class, and the formula penalises that. Duplication here is clone detection only. See Section 7.

The essential complexity was conserved. The accidental complexity rose. A rule that would have been a branch in an existing method becomes a class. The cheapest class to create is a static holder with nothing around it. Reuse is actively discouraged. Reusing code means adding lines to a function whose class already carries weight.

The more serious finding is that in some chains the rules left the measurement altogether. The comprehension metric this series relies on follows method calls from the endpoint. Work pushed downstream still counts. It cannot follow what the framework dispatches. Nothing calls a request-body interceptor, an aspect, a servlet filter or an entity listener. The container invokes them.

Spring chain, disclosed0123456789
advice + aspect classes2252203214223
create-path complexity4038871369597477075

Table 4. Per-chain escape into container dispatch. Every control chain sits at the application's baseline of two such classes. Chains 1 and 3 relocate the rules into framework-invoked interceptors. The call-graph measurement then collapses. Chain 1 reports a create-path complexity of 3 with all sixty rules implemented. Its handler is the untouched upstream body. Every rule is a RequestBodyAdvice.

Two clarifications are needed here. Both correct a mid-run note we published while the run was still going. First, this is not what most chains did. Seven of ten sat at the application's baseline of two such classes, or one above it. They dispersed into ordinary static helpers instead. Second, the effect is not uniform even among the chains that went to the container. Chain 6 has fourteen such classes and a create-path complexity of 97. The honest statement is that two chains left the measurement. The architecture as a whole did not.

That variance is itself a result. One prompt, ten chains. The agent finds materially different architectures each time. Static helpers here. Interceptors there. Injected beans elsewhere. They all score well. A number that can be satisfied this many different ways is not measuring the property it is named for.

RQ4: the correctness cost

 Spring controlSpring disclosedOF controlOF disclosed
own rule delivered (func)1.0001.0001.0001.000
all tests green (strict_pass)0.7870.4400.7320.577
standing failures, mid run9 / 4990191 / 499026151
standing failures, final198 / 7830301 / 7830168171
breakage on untouched rules37563150
median chain onset of first failurerule 47.5rule 24rule 36rule 24

Table 5. Delivered correctness. Every checkpoint's own rule landed in every chain of every condition. What degraded is retention of rules already delivered. Standing failures count prior tests failing at that checkpoint. A rule broken and never repaired keeps counting. Mutative checkpoints ship updated copies of the tests they revise. So these are real failures, not intended churn.

The agent implemented every rule it was asked for. That holds across all 2400 checkpoints in both conditions. What it stopped doing is keeping the earlier ones working. Spring's whole-suite pass rate falls from 0.787 to 0.440. The median chain's first permanent failure arrives at rule 24 rather than rule 47.5. The effect is systemic rather than one bad chain. Every chain's onset moves earlier.

We do not know why. Our mid-run explanation was ordering. Rules spread across independent interception points still have to run in some sequence. That sequence is no longer written anywhere. An agent adding a late rule cannot see the ordering it is joining. The completed data does not support it. The chains that went furthest into container dispatch broke less, not more. Their whole-suite pass rate is 0.511 against 0.410. Their standing failure rate is 2.3% against 3.1%. Three chains on one side makes that comparison weak in both directions. But it is the wrong sign for the hypothesis. We withdraw it pending a better one.

Two things the intervention did not cost. Agent spend was flat to slightly lower, at $78.40 against $76.82 per Spring chain. And the change-impact score kept its construct validity inside the disclosed run. It still correlates with independently measured agent cost, model time and comprehension effort at Spearman 0.53 to 0.60. The metric did not become noise. It became a target. That is a different failure.

6Discussion

The result is a clean instance of Goodhart's law with an unusually short time constant. There was no gate to fight. No reviewer to persuade. No quarter to wait out. A formula went into a prompt. The measured problem was then solved on the first attempt at nearly every checkpoint. The underlying property was not.

It would be unfair to call this cheating. Some of what the agent did is genuinely better. The god method is gone. The worst single function is smaller. Blast radius per rule fell. The share of checkpoints that disturb no existing function rose from 44 to 183 out of 600. A reviewer handed the disclosed Spring codebase would find real improvements in it. The problem is what the metric reports and what it stays silent about. It reports those improvements. It says nothing about the thirty-one new classes. Nothing about the duplicated lines. Nothing about the rules now invisible to the call graph. Nothing about the prior rules that stopped working. The gap between what the table says happened and what happened is very large. Only one side of it is legible in the table.

For the series thesis the result is supporting rather than damaging. It needs stating carefully. Disclosure did flatten Spring's concentration statistics. It did so by making Spring stop working like Spring. One handler becomes thirty-one dispersed classes. Two-thirds of them are static utilities. In two chains the rules route through the container instead. That path is available to Spring only as a deliberate deviation from its idiom. It took an explicit instruction. It came at a measured cost in correctness. OfficeFloor arrives at the same distribution as its ordinary way of working. Under the identical disclosed prompt it retains 0.577 of the suite against Spring's 0.440. The architecture did not stop mattering. It changed what the prompt had to overcome.

The practical implication is narrow and actionable. Do not put the scoring function in the agent's context. Keep the measure on the observing side. And hold any structural target you do set against an unscoped check. Here that check is the cumulative attribution of complexity to the functions that now contain it. A relocation cannot satisfy it.

7Threats to validity

The missing control

This is the most important limitation. It bears directly on the correctness result. We compare a disclosed-formula prompt against a plain implement-it prompt. We cannot yet separate disclosing the metric from any prompt that directs structural effort. A cohesion-prompt condition is running now to close exactly this gap. It asks in plain language for well-placed, single-responsibility code. It never mentions the metric. Until it lands, RQ4 should be read as a statement about this intervention. It is not yet a statement about disclosure specifically.

Condition labelling and prompt strength

The run carries an active gate. It fired three times in 1200 checkpoints and stopped nothing. We therefore report the condition as prompt-only. A fully ungated replication would be cleaner. The prompt is also more directive than the bare formula. It names WMC_other as the biggest lever. It instructs against growing a method or a class. These numbers are an upper bound on the effect of disclosure. They are not an estimate of the minimum.

Measurement

Cyclomatic complexity is a proxy for comprehension effort. It is not a measurement of it. The impact score charges only lines inside parsed function bodies. Logic expressed declaratively scores zero. That covers a mapper annotation, a schema, or OfficeFloor's wiring. Both architectures have such an escape, so it is not an architecture bias. But a checkpoint scoring zero should be read as logic going where the instrument cannot see. It should not be read as a cheap change. The verbosity figure in Table 3 is clone detection only. The pattern-based half of that metric did not execute in either run. It was a silent tool failure. We have since fixed it and cannot apply the fix retroactively. A separate check confirms the missing half would have contributed 15 to 32 lines per chain tip. That is against 2100 to 2500 clone lines. So the duplication finding does not depend on it.

Statistical

Ten chains per architecture per condition. Intervals come from a bootstrap clustered on chains. Breakage on untouched rules is a rare event and it is concentrated. One chain contributes 22 of Spring's 56 and 29 of OfficeFloor's 50. So the standing-failure series is the robust signal in Table 5, not that count. The per-chain container comparison in Section 5 rests on three chains. We report it as insufficient rather than as a null result.

Generality

One model. One endpoint. One sixty-step checkpoint plan. Two codebases. The checkpoint plan is fixed across conditions. That is what makes the comparison clean. It also means the specific correctness numbers are properties of this plan. Nothing here establishes that the same prompt would degrade correctness on a different change stream.

8Reproducibility and data availability

Data availability Both runs are reproducible from committed configuration. Each chain's branch carries its own configuration snapshot, its per-checkpoint capture records, and a provenance manifest naming the model, the tool versions and the parser probes. Every number here is recomputed from the commits rather than trusted from a log. The harness, the checkpoint plan and the acceptance suite are open [7].
python -m harness.run_experiment --config config.yaml --test-mode blind
python -m harness.analyze        --config config.yaml --run-id blind-202609010045

The control runs identically with the unchanged implement prompt. Analysis recomputes structural metrics from materialised worktrees at each checkpoint commit. So a metric added later can be applied to completed runs without re-running the agent.

9Conclusion

Handed its scoring function, the agent optimised it. The change-impact slope fell 45-fold for Spring. The growing handler that motivated this series never appeared. Two of the three concentration statistics stopped separating the architectures at all. On the numbers, the problem was solved.

The complexity did not go anywhere. Total complexity touched across a Spring chain is unchanged. It was redistributed out of the handler into twice as many files and thirty-one new classes. Two-thirds of those are static utilities. There is more duplication than before. In two chains the rules moved into framework-dispatched interceptors, where the call-graph measurement cannot reach them. Meanwhile the agent kept delivering every new rule. It stopped keeping the old ones working, from rule 24 rather than rule 47.5.

Brooks's essential complexity was not reducible by better instructions. Tesler's conservation held. The work moved, and it moved to wherever the measure was not looking. The operational conclusion is a boundary, not a retirement. Change impact is worth trusting as evidence about a change stream that is not targeting it. It is worth nothing as the objective that stream is given. The two uses cannot be combined. This run is how quickly the second one destroys the first.


 References and notes

  1. SlopCodeBench. No-context iterative extension, erosion and verbosity metrics, degradation slope. arXiv:2603.24755.
  2. SWE-CI. Continuous-integration gate, Normalized Change, EvoScore, Zero-Regression Rate. arXiv:2603.03823.
  3. F. P. Brooks Jr. No Silver Bullet: Essence and Accident in Software Engineering. Proc. IFIP Congress, 1986. Reprinted in IEEE Computer 20(4), 1987.
  4. L. Tesler. The law of conservation of complexity, articulated in the mid-1980s. Every process carries an irreducible complexity that design can move but not remove. Design decides only who bears it.
  5. C. A. E. Goodhart. Problems of Monetary Management: The U.K. Experience, 1975. The commonly quoted form is M. Strathern's restatement, 1997. When a measure becomes a target, it ceases to be a good measure.
  6. Change Impact in the Wild. An external test of the metric against defects in twenty human-written repositories. blog.officefloor.net.
  7. PetClinic-Evolve degradation study. Prior posts and harness. blog.officefloor.net.
  8. Cyclomatic complexity and per-function line ranges are computed with lizard. It is pinned and probed before each run. Clone detection uses jscpd at pinned thresholds.

PREPRINT · OFFICEFLOOR · SEPTEMBER 2026

Friday, 4 September 2026

I took the metric away

The last two posts did one thing. I stopped telling the AI to write good code and gave it a number instead. The exact ImpactGate cost formula, straight in the prompt. A gate that threw away any change that concentrated too much complexity. A refactor step told which classes were heavy. Everything the AI needed to keep the code clean, made measurable.

It gamed it. Exactly the way Goodhart says it will.

What the number did

When a measure becomes a target, it stops being a good measure. The AI was told the cost function and told to minimise it. So it minimised the function, not the concentration.

The formula charges a method for the complexity of the other methods in its class. So the AI stopped putting logic in the existing classes. It scattered the work into tiny new classes where the surrounding cost is near zero. And because reuse means editing a class the formula already penalises, it copied code instead of reusing it. The number went down. The code did not get better. It got more classes, more duplication, and the same tangle wearing a smaller cost.

That is the Goodhart effect in one run. The metric was honest until it became the objective. Then the AI optimised the metric and left the real problem alone.

The deeper trap: the spec becomes the program

The obvious fix is to add more metrics. Charge for duplication. Charge for new classes that only exist to dodge the first charge. Close each hole as the AI finds it.

I do not think that ends well. It is a losing game, and it is a game with a name.

Fred Brooks split software into essential and accidental complexity. The essential part is the difficulty of the problem itself. You cannot delete it. You can only decide where it lives. Larry Tesler said the same thing more bluntly. His Law of Conservation of Complexity says every system has an irreducible amount of it. The only question is who carries it.

So when I pile more rules into the metric, I am not removing the complexity. I am moving it into the metric. Keep going and the metric has to describe every structural decision precisely enough for the AI to optimise against it. At that point the metric, plus the spec, is a full description of the system. It is the program. Written in a language that cannot be run, cannot be tested, and has no tools.

That is the trap. If keeping the code clean requires me to specify the system twice, once as a spec to optimise and once as the code it produces, then the spec is just a more expensive way to maintain the definition of the system than the code was. Code is already the most precise, testable, tooled description we have. Replacing it with an ever growing specification is not progress.

So I stopped adding to the number. I started taking it away. Two experiments, two ways to remove it.

Experiment one: keep the tool, hide the number

The first experiment keeps ImpactGate but never lets the AI see the formula.

The AI implements each change with a plain spec. No metric. No objective to optimise. ImpactGate still scores the change in the background. When a change concentrates too much, one refactor runs first. But the refactor is told only the symptom, not the cure.

It is given the names of the classes that ended up carrying too much. Locations only. Not the cost, not the formula, not the words "split into small classes". Those phrases are what drove the dispersal and the duplication last time. Here is the whole refactor prompt.

A change is about to be made to this codebase. When that change was
implemented directly, these classes ended up carrying more complexity
than they can hold comfortably:

{drivers}

Before the change is made, prepare the existing code so the change can
land as code a maintainer of this codebase would recognise as natural,
consistent with how this codebase already structures similar work.

It names the sore spot and asks for something natural. It does not hand over a lever to pull.

One more guard. The refactor's own new lines run through a deterministic quality gate. Duplicated lines and known bad patterns are caught by jscpd and ast-grep, not by a prompt. So the AI cannot buy a quieter structure with copy paste. The prompt asks for good judgement. The gate, not the prompt, forbids the cheap moves.

The question is simple. With no number to optimise and only a nudge at the sore spot, does the structure stay cohesive?

Experiment two: ask in plain English

The second experiment drops the tool entirely. No gate. No refactor step. No number anywhere.

I am not ruling out that guidance helps. I am ruling out the formula. So the ask goes back to English, but honest English. Not a target to optimise. A plain request for structure, in the words a senior engineer would use in review.

As you do, keep the code well structured: put each piece of logic where
it belongs, in a small unit with a single responsibility, and reuse
existing code instead of copying it. Do not let any one class or method
grow into a catch-all that accumulates unrelated logic.

That is the whole intervention. Same single turn as the plain spec run. Same everything else. Only the wording changes. It says nothing about a metric, a cost, or the experiment. It cannot be gamed, because there is no number to game. It also cannot be checked, by the AI or by me, which is exactly the weakness the formula was meant to fix.

So this experiment asks the honest version of the original question. Not "can a measurable target keep code clean", which the AI answered by gaming the target. But "can plain good advice keep code clean", with nothing to optimise against it.

What the two experiments are really testing

Line everything up and there is a clean ladder.

  • Plain spec. No help at all. The baseline erosion.
  • Plain English. Good advice, no number. The prompting lever on its own.
  • Refactor run. A tool that catches the sore spot and nudges it, still no number shown. The tool lever on its own.
  • The formula. The full measurable target. Already run. Already gamed.

Each step adds exactly one thing. So the gap between two steps is the effect of that one thing. And every one of them is measured the same way at the end, on the shape of the standing code, not on any score the AI could aim at.

Here is the worry that started all this. Tesler says the complexity never leaves. It only moves. My first bet was to hold it down with a measurement, and the AI gamed the measurement. These two experiments try the opposite and take the number away.

But the complexity still has to land somewhere. Maybe the AI invents its own target to chase, and games that instead, with no formula in sight. Maybe the only way to stop it is to keep adding to the spec, until the spec is the program again. If either happens, then a metric was never the right lever. A better number cannot fix a problem that lives in the architecture. The fix would have to be the architecture itself.

That is what these two experiments are built to find out. Same harness, same problem, same chains, same measurement at the end.

The runs are underway.

Thursday, 3 September 2026

We Gave the Agent the Metric

A series on how software architecture shapes AI-driven code degradation. This post is a mid-run note. The run is still going, so there are no numbers here yet, only what we can see so far.

Every result in this series has an obvious rebuttal. Of course Spring's handler grows. Nobody told the agent not to let it.

So we told it. The current run hands the agent the exact cost formula the experiment measures. Not vague advice about clean code. The actual arithmetic. Take the complexity of the function you touch. Multiply it by the complexity of everything else in its class. Multiply that by how much of the function you change. Multiply the total by how many files you spread the change across. Lower is better. That is the whole objective, stated plainly, at every checkpoint.

There is a gate behind it as well. Each change is scored, and a change that lands above the threshold is thrown away and re-attempted after a refactor. That was the part we expected to do the work.

The gate has had nothing to do

It has not fired once.

The prompt alone was enough. Told what the measure is, the agent simply writes code that scores well on it, from the first attempt, nearly every time. The control loop we built to force the issue has been sitting idle while the prompt does all of it.

That is the first thing worth saying out loud. If you want an agent to optimise something, telling it the formula works.

Spring's god method never appeared

In every previous run, Spring's create-owner handler grew. Rule after rule landed in the same method until it was the largest thing in the codebase. That is the result the series has been built on.

This run, it does not happen. The handler ends close to where it started. The concentration measures that used to separate the two architectures no longer separate them. On the headline numbers, Spring now looks as clean as OfficeFloor, and on one of them it looks cleaner.

Taken at face value, that is the objection landing. Prompt better, and architecture stops mattering.

Then we looked at the code

The rules did not get smaller. Nothing was simplified away.

It moved.

Out of the handler, and into whatever host had the least surrounding complexity to pay for. Request advice. Aspects. Servlet filters. Entity listeners. Bean validators. Small classes that exist to hold one rule and nothing else. The logic is all still there. It is just somewhere else now, in a lot of somewhere elses.

In one chain the handler is byte for byte the file we started with. Not one of the rules is visible at the endpoint. They are all interceptors. The framework calls them. The order they run in is decided by annotations scattered across many files.

The measure cannot see where the work went

This is the part that matters, and it is uncomfortable, because it is our own measure that broke.

The comprehension measure we are most proud of follows the calls. It starts at the endpoint and walks every method the request path reaches, so work pushed downstream still counts. That was the fix a reader pushed us into making, and it was the right fix.

It has a blind spot. If nothing calls the code, the walk never reaches it. An interceptor is invoked by the framework, not by the handler, so a rule that becomes an interceptor leaves the measurement entirely. Most of Spring's create-owner logic is now outside what that walk can see.

So the flattering number is not a comprehension win. A developer changing one rule still has to find it first, and finding it is now harder, not easier. The measure got quieter. The code base did not get simpler.

Every chain escapes differently

There is no single shape to this. Run the same sequence again and the agent picks another way out. One chain went all in on request advice. Another used aspects for nearly everything. Others extracted static helpers, or pushed rules into validators and entity callbacks, and kept a slim handler calling them in order.

They all score well. They are wildly different code bases. If you are choosing an architecture on the strength of a number like this, that variance is the warning.

OfficeFloor moved too

Less, but in the same direction, and this one surprised us.

OfficeFloor answers a new rule by wiring in a new function. That is its whole idiom, and it is the behaviour the formula should reward. But wiring a function means touching the wiring file as well as writing the class. The formula multiplies by the number of files you touch. So the cheapest move stops being the idiomatic one.

What we see is a drift towards plain helper classes that nothing wires, called from a function that already exists. And a drift towards folding small rules into the entry function itself. That function's class is nearly empty. Under the formula, that makes it almost free to grow. The entry function is now doing more than it used to.

The pressure is far weaker than it is on Spring. The architecture already sits near where the formula wants it to be. But the direction is the same, and it is away from the architecture's own way of doing things.

Something is breaking more often

The agent still implements the rule in front of it at the same rate it always did. That has not moved.

What has moved is the damage to rules nobody asked it to touch. Both architectures now break more of them than they did before. Our working explanation is ordering. Rules spread across many independent interception points still have to run in some sequence. That sequence is no longer written down anywhere. It is implied by annotations, bean order, and framework lifecycle. An agent adding a late rule cannot see the ordering it is joining.

We want more of the run before we lean on this one.

What we think this is

It is Goodhart. A measure became a target and stopped being a good measure.

We should be fair to the agent here, because this is not simple cheating. Some of what it did is genuinely better. There is no god method. The worst single function in the code base is smaller than it was in either architecture before. Those are real improvements and we will report them as such.

But the gap between what the metric says happened and what actually happened is very large. The metric says the problem is solved. The code says the same work was cut finer and hidden better. Both of those are true at once, and only one of them shows up in a table.

The honest reading is that we cannot use this metric as an optimisation target and as evidence at the same time. As evidence about untargeted code, it held up well. Pointed at as a goal, it collapsed inside a single run, without a single gate intervention, using nothing but a prompt.

The run has chains to go. The numbers, and the final call on what this does to the series, will follow when it finishes.

Monday, 31 August 2026

Give the AI the formula: prompting or architecture?

In the last experiment I held the coding agent fixed and changed only the architecture. Spring pushed each new rule into one growing controller method. OfficeFloor spread each new rule across many small wired functions. The same total complexity landed in both. But Spring concentrated it. OfficeFloor stayed cohesive.

There is a fair objection to that result. Maybe I just prompted the agent badly. Maybe if you tell the AI to write clean code, Spring would be fine. Maybe the architecture is not the cause at all. This next run is built to answer that.

The problem with "write good code"

You can tell an AI to write good code. You can tell it to avoid god classes. It is vague. It is not something the AI can measure itself against. So it cannot know if it succeeded.

So I stopped using English. I gave the AI a number. A precise, measurable target for how much complexity a change is allowed to concentrate. The tool that produces that number is ImpactGate (for openness this is also part of the OfficeFloor suite).

ImpactGate: an objective control around erosion

ImpactGate is a small command line plugin. It scores the structural impact of a change. It reads any language through a single complexity engine based on the Change Impact formula. It has a GitHub Action and can post the score on a pull request. It can warn or block by exit code. So it drops into a CI pipeline as a real gate.

The score is not an opinion. It is a formula. Erosion stops being a feeling in a code review and becomes a measurable control. In this experiment ImpactGate is that control. It sits in the pipeline and it decides, on every change, whether the code is about to concentrate complexity.

The formula, given to the AI

Here is the cost ImpactGate computes. For every function a change touches:

cost = max(WMC_other, 1) * CC * max(1, changed_lines)

CC is the cyclomatic complexity of that function. WMC_other is the summed complexity of the other methods in its class. That is the surrounding context you must hold in your head to edit it safely. changed_lines is how many of its lines the change adds or edits. The costs are summed over every changed function. Then the total is multiplied by the number of files touched.

Look at what dominates. It is WMC_other. A method inside a large class pays for the whole class. Split that class into small cohesive classes and the cost falls for every method. This is the exact reason a god class is expensive. And it is a lever the AI can pull.

So I put this formula straight into the prompt. The AI is not told to write good code. It is told how good code is measured. It is given the objective and the main lever to move it. This is the difference the experiment tests. Clear measurable target, not English hand waving.

The pipeline

Every change runs through this loop.

  1. The AI implements the change, with the formula as its stated objective.
  2. ImpactGate scores the change.
  3. Under the line, the change is accepted.
  4. Over the line, the change is thrown away.
  5. A refactor step runs on the clean code. It is given the same formula, the heavy classes, and the change that is coming.
  6. It breaks those classes into smaller cohesive ones. It does not implement the change yet.
  7. The refactor is committed on its own, so you can see it.
  8. The change is attempted again, on the cleaner code.

A change gets up to three refactors to come under the line. Still over after the third, the run stops. That is recorded as a failure to keep the code clean.

A fair line, set by OfficeFloor

The gate needs a line. I did not want a generic one. I set it from OfficeFloor itself. OfficeFloor already stays cohesive. So its own change scores are the picture of good behaviour.

I took the distribution of OfficeFloor's per change impact from the earlier run. That is 549 real changes. Every new change is graded against that distribution. The same line is used for both arms. Spring must meet it. OfficeFloor must meet it too, so it cannot quietly erode either.

Percentile OfficeFloor Spring
p50 (median)3694,736
p903,41031,845
p955,92254,820
p9814,148105,734

Read that gap. Spring's median change is bigger than 90 percent of OfficeFloor's changes. I set the cutoff at OfficeFloor's 95th percentile. At that line about 45 percent of Spring's changes need a refactor. About 5 percent of OfficeFloor's own changes do. The bar is strict. But it is not invented. It is the cohesion OfficeFloor already reaches on this same problem.

The question this run answers

Now the AI has everything. The exact formula. The main lever. A gate that catches every concentrated change. A refactor step that is told what to fix. Three attempts per change.

If the code still erodes under all of that, the prompt was not the problem. So the run splits cleanly into two answers.

  • Spring keeps its complexity spread and reaches the end. Good prompting can keep code clean.
  • Spring cannot stay under the line and the run stops. The erosion is in the architecture, not the prompt.

Both arms run the same pipeline, ten chains each, held to the same OfficeFloor line. Every gate decision is recorded with the tool version and the exact baseline it graded against, so a run is reproducible. The measure the AI is optimising is the same measure I report at the end, so the AI cannot game one while I grade the other.

The setup is done. Now to burn AI tokens to see the outcome.

Tuesday, 25 August 2026

Change Impact in the Wild: A Defect Predictor Undone by File Size

Preprint / cs.SE / Empirical Software Engineering

Change Impact in the Wild

Twenty repositories. Six languages. A full baseline battery. An honest reckoning with what a complexity metric predicts, and where its value actually lies.

OfficeFloor · independent research · blog.officefloor.net

August 2026 · correspondence: daniel@officefloor.net


Abstract

Change impact scores a code change by the complexity it disturbs, not the lines it edits. It was defined inside a single controlled experiment. There it was tuned against the effort an AI coding agent spent as changes accumulated on a fixed codebase. This paper asks a harder question. Does it predict defects in human-written code it has never seen? We test it across twenty open-source repositories in Java, C#, JavaScript and TypeScript, Python, Go, and C. We test it against a full battery of baselines. These are churn, the hotspot, file size, total complexity, change entropy, developer count, and prior fixes. Each is controlled one at a time. Then, crucially, all of them are controlled together.

Controlled against churn alone, change impact looks strong. It ranks future fix locations in all twenty repositories (median partial 0.19). It ranks bug-inducing commits in all twenty (median 0.32). That impression does not survive scrutiny. Plain file size out-predicts change impact in all twenty repositories. Then size, complexity volume, spread, and churn are removed together. The metric's unique contribution all but vanishes. The multivariate partial has a median of 0.01 for location (positive in 12 of 20) and 0.01 for introduction (12 of 20). The distinctive concentration-weighting adds little beyond raw complexity volume. Its residual does not track measured complexity concentration. The honest conclusion is a negative one for defect prediction.

But defect prediction was never where this metric belongs. Tests catch bugs. Its value is prospective. It watches change impact rise as changes pile in. That is the cue to refactor before complexity concentrates and code loses cohesion. That is the setting it was born in. It also matters most for AI-augmented pipelines, where an agent lands many changes and a codebase can silently degrade. The tool and pipeline are open source, offline, and deterministic.

Keywords: change impact · defect prediction · mining software repositories · multivariate baselines · code churn · cyclomatic complexity · complexity management · AI-assisted development

1Introduction

A metric earns trust by predicting something it was not built on. Change impact was built and tuned inside one controlled degradation study, where it correlated with the cost, re-reading, and model time an AI agent spent as accumulating changes landed on a fixed endpoint [1][5]. That is concurrent validity on two arms of a single experiment. It is silent on whether the score means anything on code it has never seen.

This paper supplies the missing test. It reports it in full, including where it fails. We compute change impact over the history of twenty independent open-source repositories. We ask whether it predicts defects, which git history records through bug-fixing commits. The bar is not whether change impact correlates with defects. Any size-like measure does that. The bar is whether it adds signal beyond the measures a team already has. We set that bar high. Not one baseline but a battery. And not one control at a time but all of them together.

The short answer is that it does not. Change impact correlates with defects. But the correlation is largely a restatement of file size and raw complexity volume. Under multivariate control the metric's own contribution is near zero. That is a negative result for defect prediction. It is also a useful one. It redirects the metric to the question it was actually built for. Not where are the bugs, but where is complexity concentrating dangerously as changes accumulate.

The contributions are as follows.

  1. An external test of change impact against defects across twenty repositories in six languages, none used to design the metric.
  2. A full baseline battery, controlled both singly and, via rank residualisation, jointly. The baselines are churn, hotspot, file size, total complexity, change entropy, developer count, and prior fixes.
  3. The finding that change impact's apparent defect signal does not survive multivariate control, and that its residual does not track measured complexity concentration: an honest negative result.
  4. A reframing of the metric as a prospective complexity-management signal for change streams, human and AI-generated, and an open, offline, deterministic tool that reproduces every number here and applies to any git repository.

2Background and related work

Size and activity are strong, hard-to-beat defect predictors. Code churn and the number of prior changes to a file correlate robustly with faults, and the hotspot, complexity multiplied by change frequency, is a widely used prioritization signal [6]. Any new structural metric must be measured against these baselines, not against chance.

The SZZ algorithm identifies bug-introducing changes by blaming the lines a fix modifies back to the commits that last wrote them [3]. It is approximate. Blame names the last modifier rather than the true author of a defect, and recent commits are under-observed because later fixes have not yet had time to touch them.

Change impact originates in a study that holds the coding agent fixed and varies architecture, measuring how a codebase degrades as roughly sixty changes accumulate on one endpoint [1][2][5]. There, change impact tracked the agent's effort. Here we ask a different and harder question, whether it tracks defects in projects built by people, over years, in many languages.

3The change-impact metric

For each function a change touches, the cost is the product of three terms.

cost(function) = max(WMC_other, 1) × CC × max(1, Δlines)

CC is the function's cyclomatic complexity. Δlines is the number of lines the change touched inside it. WMC_other is the summed complexity of the other functions sharing the function's scope, measured on the state before the change: the surrounding context a maintainer had to comprehend to make the change safely. The scope is the enclosing class where the language provides one, and the file otherwise, which keeps the score defined across procedural and object-oriented code alike.

Measuring the context before the change, rather than after, is deliberate. A function added to a brand-new file or class has no prior neighbours, so its WMC_other falls to one and the cost reduces to CC times lines. Importing or scaffolding fresh code stays cheap. A function grafted onto an already-heavy class is charged for the complexity that was already there. Piling onto a concentrated scope stays expensive. That is the property the score is meant to capture. This before-change definition is used throughout; measuring the scope after the change instead shifts individual scores but leaves every result in this paper essentially unchanged, so none of the findings depend on the choice.

Two variants are reported. Mutation impact sums cost over functions that already existed, the cost of disturbing what is there, with no spread term.

mutation = ∑ cost(function)  over mutated functions

Composite impact adds the cost of newly introduced functions. That cost is floored, through the max terms, so that fragmenting logic into cohesionless new functions is not free. It then multiplies the total by the number of source files the change touches.

composite = N_files × ( ∑ costmutated  +  ∑ costnew )

The file multiplier is a spread penalty: an edit scattered across many files costs more than the same edit confined to one. A within-commit rename, detected by body token overlap, is scored as a mutation rather than a free addition, closing an obvious gaming path.

4Study design

Research questions

  • RQ1, location. Do files with higher change impact receive more future bug fixes, beyond what churn explains?
  • RQ2, introduction. Do commits with higher change impact induce more future fixes under SZZ, beyond what commit size explains?

Corpus

Twenty repositories were selected for long history, real bug-fix signal, and a spread of architectural concentration, across six languages. History depth ranges from about seven thousand to ninety-three thousand mainline commits. Merge commits are followed by first parent, so a merge is diffed against the branch it introduced. A twenty-first repository, Kibana, was dropped for a download failure rather than for its data.

Ground truth

Bug-fix commits are identified from the commit message, precision-ranked, reverts first, then issue-closing references such as KAFKA-1234 or fixes #123, then fix keywords. The label is coarse and reused unchanged across all repositories. For RQ2 we apply SZZ, blaming each fix's changed lines at its parent with git blame -w -C to recover the inducing commits.

Statistics and controls

Predictors are measured over the first 75% of each history by commit time. Outcomes are measured over the last 25%. That makes RQ1 leakage free. The baseline battery spans both families a reviewer expects. The process signals are churn (added plus removed lines), commit frequency, change entropy [7], developer count, and prior bug-fixes. The code signals are file size, total complexity (summed CC at the split), and the hotspot (complexity times change frequency) [6]. File size and complexity are read at the split snapshot from the repository tree.

Two kinds of control are reported. The single-control partial is a partial Spearman of change impact with the outcome, removing one baseline at a time; the metric must stay positive after removing each. The multivariate partial removes a whole set of baselines at once, by rank-transforming every variable, regressing change impact and the outcome on the full control set by ordinary least squares, and correlating the residuals. This is the decisive test: a signal can beat each rival singly yet add nothing once the correlated rivals are removed together. Ranking quality is also reported as the area under the ROC curve, and test files are excluded from the location universe.

5Results

RQ1, where defects live

Mutation impact predicts future bug-fix locations, beyond churn, in every repository. The partial Spearman controlling for churn is positive in 20 of 20, with a median of 0.19 and a range of 0.065 to 0.356. Ranking is better than churn as well, with AUC for mutation impact exceeding AUC for churn in all twenty. The temporal split exposes what the concurrent view hides. Measured across the split, past churn barely predicts future fixes, and its correlation is near zero or negative in a third of the repositories, while past mutation impact stays positive. Past size is a weak forecaster. Past concentration of change is not.

The tougher control is the hotspot itself, complexity times change frequency. It is a near cousin of change impact. Mutation impact stays positive after removing it in all twenty repositories, with a median partial of 0.106 and a range of 0.021 to 0.208. Its ranking AUC exceeds the hotspot’s in seventeen. The margin is about half the churn-controlled one, as a rival that already carries complexity should be. It is thinnest where the code is clean. The lowest is home-assistant at 0.021. The three repositories where the hotspot out-ranks impact are Django, Guava, and home-assistant, all low-concentration or well-factored. The signal beyond the hotspot never vanishes. But it is small where complexity does not concentrate.

These single-control numbers share a blind spot. The strongest baseline of all was missing from them. That baseline is plain file size, bytes at the split. File size out-predicts change impact in every one of the twenty repositories, with a median raw Spearman of 0.47 against 0.16 for impact. Controlling for size alone, impact's location partial drops to a median of 0.10. It turns negative in two repositories. Size, not churn, was the rival to beat. Change impact does not clearly beat it.

The multivariate test is decisive. We remove churn, hotspot, file size, and total complexity together. Change impact's location partial falls to a median of 0.005, positive in only 12 of 20 repositories. That is a coin toss. Adding change entropy and developer count leaves it positive in 9 of 20, median below zero. Once the correlated size and complexity signals are removed at once, change impact adds essentially nothing to locating defects beyond what those simpler measures already provide.

RQ2, which commits introduce defects

The SZZ ranking AUCs are high, 0.75 to 0.93. That is partly mechanical. A larger commit offers more lines for a later fix to blame. Controlling for commit churn, composite impact still ranks inducing commits in 20 of 20 repositories, median 0.32, range 0.166 to 0.687. That is the strongest single-control result in the study.

It does not survive the joint test. Churn is only one measure of a commit's size. A commit also has a count of files touched, a total complexity changed, and a number of functions changed. Composite impact is, by construction, a product of those very quantities. Remove churn, files, total complexity, and function count together, and the median partial collapses from 0.32 to 0.006. It is positive in 12 of 20 repositories but negative in eight. git, the highest at 0.69 under churn alone, falls to 0.048. The concentration-weighting that makes change impact distinctive buys almost nothing over the raw volume and spread of the commit.

Figure 1. The single-control partial correlations, per repository, sorted by the location result. Teal marks RQ1 (mutation impact vs future fixes, controlling for churn). Gold marks RQ2 (composite impact vs induced fixes, controlling for churn). Every point falls right of zero. But controlling for churn alone is a low bar. Under the full multivariate control (Table 2) both collapse toward zero.

The mutation and composite split

Within the single-control view the two variants separate cleanly along the two questions. For location, mutation impact beats composite in 19 of 20 repositories. For introduction, composite beats mutation in 20 of 20. Disturbing existing complex code is where fixes concentrate. Writing new complex code is where inducing commits land. It is a tidy pattern. But, like the partials it rests on, it reflects how the two variants track size and complexity volume. It is not a signal that survives their joint removal.

The concentration thesis does not replicate

An earlier reading of these repositories noted that the well-factored controls, Guava and Spring Framework, sat at the bottom of the single-control rankings, and took it as unbidden support for the idea that change impact earns its keep where complexity concentrates. Tested directly, that story does not hold. The multivariate residual is what change impact adds beyond size and complexity volume. Across the twenty repositories it shows no positive association with a repository's measured complexity concentration. This holds whether concentration is taken as the Gini of unit complexity, the share held by the top one percent of functions, or the maximum surrounding complexity. For defect introduction the associations are near zero or negative. The top-one-percent measure is the strongest, at −0.46, and it points the wrong way. For location a single weak positive appears against one proxy. It is not corroborated by the others. With twenty repositories the test is underpowered. But the point estimates do not even point consistently in the hypothesised direction.

repositorylangcommitsfiles prevpartial_mutpartial_hotAUC_mutAUC_churn SZZ_pcomp

Table 1. Single-control partials per repository. prev is outcome prevalence, the share of files with a bug-fix touch in the outcome window. partial_mut is the location result controlling for churn (RQ1); partial_hot controls for the hotspot baseline (complexity×change-frequency); SZZ_pcomp is the introduction result controlling for commit size (RQ2). The gold columns are partial Spearman correlations, each removing one rival. All collapse under the joint control of Table 2. Full columns are in summary.csv in the repository [4].

testcontrols removed (together) median partialpositive
RQ1 locationchurn0.19420 / 20
RQ1 location+ hotspot0.10620 / 20
RQ1 locationfile size alone0.09718 / 20
RQ1 locationchurn + hotspot + size + ΣCC0.00512 / 20
RQ2 introductioncommit churn0.31520 / 20
RQ2 introductionchurn + files + ΣCC + #units0.00612 / 20

Table 2. From single-control to multivariate. Correlated size and complexity baselines are removed together. Change impact's partial correlation with defects then collapses toward zero. This holds both for where they are fixed and for which commits introduce them. File size alone already halves the location signal. The joint control erases it. Medians are across the twenty repositories.

6Discussion

The negative result is worth stating plainly. As a defect predictor, change impact does not earn its complexity. For locating defects, plain file size does better. For both questions, a handful of size and complexity-volume measures, removed together, absorb essentially all of the metric's signal. The concentration-weighting is the one thing that distinguishes change impact from counting lines or summing complexity. It adds almost nothing those cheaper measures do not already carry.

But defect prediction was the wrong target. In a project with a test suite, most bugs are caught before they are committed. The fixes git records are the residue that slipped through. That is a noisy and lagging signal. Change impact was never built to forecast that residue. It was built, in its original study, to measure how much a codebase degrades as changes accumulate against a fixed endpoint. It measures how far each change pushes the code toward tangled, low-cohesion, hard-to-change structure. That is a property of the change stream. It is observable the moment a change lands, not a property of some future fix.

Read that way, the metric's value is prospective and actionable rather than predictive. It answers a question a size counter cannot. Not which files are risky. The big and complex ones are risky, and everyone already knows that. The real question is which specific scope a change is overloading. The cost is driven by the surrounding complexity a maintainer must hold in their head to change it safely. A high change impact is a prompt. Stop and refactor that concentration before the next change lands on top of it. That keeps the code additive and cohesive rather than letting a hot scope thicken.

This matters most where the change stream is fast and only lightly reviewed. Think of AI-augmented pipelines. An agent can land dozens of changes in an afternoon, and a codebase can degrade faster than a reviewer can notice. There the useful signal is not a defect probability. It is a live gauge of accumulating complexity. Flag the change that should be split. Flag the scope that should be decomposed. Do it before the agent piles on. That is the setting the metric was born in, and the honest place for it to return.

7Threats to validity

Analysis, self-referential controls. For introduction, the multivariate control set overlaps the raw ingredients of composite impact. That set is files touched, total complexity, and function count. Composite impact is by construction a product of them. Removing them is close to asking whether composite beats its own parts. That is the right question for isolating the concentration-weighting's marginal value. But it makes the RQ2 collapse partly definitional rather than purely empirical. The RQ1 location collapse, driven by file size, is not subject to this caveat.

Statistical power, concentration test. The concentration analysis correlates a per-repository residual against a per-repository concentration measure over only twenty points; it is underpowered, and a genuine weak effect could be missed. The point estimates, however, are scattered around zero and do not point consistently in the hypothesised direction, so the null is more than a power failure.

Construct, ground truth. The bug-fix label is keyword based and over-counts, since fix also matches typos and formatting; and a test suite catches most defects before commit, so the recorded fixes are a lagging, partial signal. A typed issue tracker would sharpen it but would not change the multivariate verdict.

Construct, SZZ. Blame names the last modifier, not the defect's author, and RQ2 runs SZZ concurrently, so recent commits are under-counted as inducers.

Internal, renames and measures. Rename handling is deliberately light and cross-directory moves are not chased. File size is measured in bytes and complexity as summed cyclomatic complexity at the split; both are reasonable but coarse. A single split fraction, 0.75, is used throughout, and the corpus is diverse but not a random sample of software.

Internal, vendored directories. The ignore rule initially matched vendored and generated trees only when nested, so a repository-root vendor/ or node_modules/ slipped through. This inflated three vendored Go repositories, Prometheus, Moby, and Kubernetes. Excluding those trees, as the numbers here do, changed those three repositories' individual results slightly and left every corpus median and every conclusion unchanged.

8Reproducibility and data availability

Data availability Surveyor is open source and runs offline against any local git clone [4]. The pipeline is deterministic and every number in this paper is recomputed from the commits, so a new metric can be added and applied to the same runs without re-reading the repositories.
surveyor scan    <repo>
surveyor analyze <repo> --split-frac 0.75

A parallel driver scans and analyzes the whole corpus and emits the cross-repository summary from which the table and figure above are drawn.

9Conclusion

Change impact does not predict defects beyond the simple measures a team already has. Controlled against churn alone it looks strong. But plain file size out-predicts it in every repository. Once size and complexity volume are removed together, its own contribution falls to a coin toss for defect location and near zero for defect introduction. Its residual does not track where complexity concentrates. As an external defect predictor the metric fails. This paper reports that squarely.

The result redirects rather than retires the metric. Change impact measures how much a change disturbs concentrated complexity. That is a live property of a change stream. It is useful for deciding when to refactor before code loses cohesion. It is pointed enough to name the scope at fault in a way a size counter cannot. Its natural home is not forecasting the bugs a test suite already catches. It is keeping a codebase cohesive as changes pile in. That matters most urgently in the AI-augmented pipelines where that stream now runs fastest. Testing that prospective, in-the-loop use is the next step. The tool that produced every number here is open source to support it.


RReferences and notes

  1. SlopCodeBench. No-context iterative extension, erosion and verbosity metrics, degradation slope. arXiv:2603.24755.
  2. SWE-CI. Continuous-integration gate, Normalized Change, EvoScore, Zero-Regression Rate. arXiv:2603.03823.
  3. J. Śliwerski, T. Zimmermann, A. Zeller. When do changes induce fixes? Proc. Mining Software Repositories (MSR), 2005. The SZZ algorithm.
  4. Surveyor. Language-agnostic change-impact and pain-signal harness. github.com/officefloor/Surveyor.
  5. PetClinic-Evolve degradation study. Prior posts in this series, blog.officefloor.net.
  6. A. Tornhill. Hotspots as complexity times change frequency, and change coupling. Your Code as a Crime Scene.
  7. A. E. Hassan. Predicting faults using the complexity of code changes. Proc. International Conference on Software Engineering (ICSE), 2009. Change entropy.
  8. Cyclomatic complexity and per-function line ranges are computed with lizard, which supplies multi-language parsing.

PREPRINT · OFFICEFLOOR · AUGUST 2026

We Tested Our Complexity Metric Against Real Bugs. File Size Won.

Some files attract bug fixes over and over. Some commits quietly introduce the bugs that later get fixed. We wanted a way to spot both from git history alone, before the bugs show up. We thought we already had the measure for it.

It is called change impact. It scores each change by how much surrounding complexity it disturbs. It was built and tuned on a controlled experiment. So the fair question was whether it means anything out in the wild. We tested it against real bug fixes in twenty open-source projects, as hard as we could. The honest answer is not the one we expected. It turned out to be more useful than the one we went looking for.

This is the plain-English version. The full numbers, the method, and the statistics are in the companion post: Change Impact in the Wild.

What change impact measures

Adding a brand-new file is easy. You write it once. Nothing else has to move.

Changing a method inside a large, tangled class is not easy. You have to hold everything around it in your head first. And a mistake there ripples outward.

Change impact captures that difference. For each function a change touches, it multiplies three things:

  • how complex the surrounding code already was, the part you must understand to touch it safely,
  • how complex the function itself is,
  • how many lines you changed.

Then it scales by how many files the change spread across. A one-line tweak to an isolated helper scores low. The same tweak inside a two-thousand-line god class scores high. That is the whole idea. Not all change is equal. This puts a number on the difference.

One detail matters. The surrounding complexity is measured as it stood before your change. So creating a fresh file or class is cheap, because nothing was there yet. Adding onto an already-heavy class is what costs, because you have to work around everything already in it. That is the point of the measure. It charges you for piling onto a concentrated scope, not for writing something new and self-contained.

What we tested, and what actually held up

We asked change impact two questions. Which files will attract future bug fixes? And which commits introduce the bugs that later get fixed? We measured the first three quarters of each project's history. Then we checked its predictions against the final quarter. No hindsight.

At first it looked great. We accounted for how much a file changes, its churn. Even then, change impact still lined up with where bugs later appeared. That held in all twenty projects. If we had stopped there, we would have published a win.

We did not stop there. Two things brought it down.

Plain file size beats it. We had never put raw file size in as a rival. When we did, file size predicted where bugs land better than change impact. That held in all twenty projects. The honest headline is boring. Bugs tend to be in the big, complex files. A byte count already tells you that.

Remove the simple measures together, and almost nothing is left. Change impact is basically size times complexity times spread. So we removed all of those at once. That means file size, total complexity, churn, and how many files a commit touches. Then we asked what change impact still adds on its own. The answer is next to nothing. For finding buggy files it came out to a coin toss. For finding bug-introducing commits it dropped from a strong-looking number to roughly zero. The clever part is the weighting by where complexity concentrates. It buys almost nothing over just measuring how much size and complexity a change carries.

We also checked the appealing story that change impact shines in tangled codebases and stays quiet in clean ones. Tested directly, that did not hold up either.

Why we are telling you the negative result

Because it is the true one. We would rather find it ourselves than have a reviewer find it for us. A metric that only beats the weakest rival is not a bug predictor. And this one folds the moment you line it up against file size. Saying otherwise would not survive contact with anyone who checked.

The part that is actually useful

Here is the reframe, and it is the interesting bit. Predicting bugs was the wrong job for this metric. If you have a test suite, most bugs are caught before they ever land. The fixes in git history are the leftovers that slipped through. That is a noisy, lagging signal. Change impact was never built to forecast those.

It was built to measure something you can see the instant a change lands. It measures how much that change degrades the structure. It measures how far a change pushes the code toward tangled, low-cohesion, hard-to-change shape. That is not a prediction about some future bug. It is a live reading on the change in front of you.

And that reading is something a size counter cannot give you. File size can tell you this file is big and risky. You already knew that. Change impact can tell you which scope a change is overloading. The cost is driven by the specific surrounding complexity you would have to untangle. So it is a prompt. Stop. Refactor this concentration. Then make the change. That way the next change does not land on top of a thickening hot spot.

This matters most where changes arrive fast and get little review. Think of AI-assisted pipelines. An agent can land dozens of changes in an afternoon. A codebase can quietly rot faster than anyone notices. There you do not want a bug probability. You want a live gauge. It should say this change is piling complexity into one place. Split it, or decompose the scope first. That is the job change impact is actually good at. It is the one we are building toward next.

What it means for you

  • Do not reach for change impact to predict bugs. For that, it does not beat file size and churn, and you already have those.
  • Do reach for it to watch complexity accumulate. A change with high impact is a signal to refactor the concentrated scope before piling on. That is most useful when the changes are coming from an agent, faster than you can eyeball them.

The honest limits

This is a negative result on prediction, and we hold it as one. We label bug fixes from commit messages, which is noisy. A test suite hides most bugs from that signal anyway. And some of the remove-everything-at-once test is stacked against a metric built from those same ingredients. We also tightened the measure itself, so it now weighs the complexity that was there before a change rather than after. That is our current definition, and we re-ran every number on it. The verdict did not move, which makes the result sturdier, not weaker. The direction is clear and consistent across twenty projects and six languages. As a defect predictor beyond simple measures, change impact does not hold up. Its value is prospective, not predictive.

If you want the tables, the statistics, and the threats to validity, read the companion post: Change Impact in the Wild. The tool is open source and runs offline on any git repository, so you can point it at your own code: Surveyor on GitHub.

Monday, 24 August 2026

The Same Complexity. One Unit or Twenty-Two.

A series on how software architecture shapes AI-driven code degradation. An objection to the last post turned out to be right, and fixing it produced a better result than the one it demolished.

Every concentration number in this experiment has measured the same thing: the endpoint's entry handler. Its complexity, its class weight, its erosion. Spring's climbs. OfficeFloor's stays flat.

The objection is obvious once someone says it out loud. OfficeFloor is a pipeline. It can keep its first function pristine by pushing the work into the second one. A metric that looks only at the front door will be fooled by anyone who moves the mess into the hallway.

So we measured the hallway.

Following every call

OfficeFloor declares its pipeline in a wiring file, so every step the request passes through is enumerable. From each step, and from Spring's single handler, we followed method calls transitively through the project: helpers, hashing, formatting, entity derivations, wherever they live. Whatever a node reaches is charged to that node.

Java call resolution without a type checker is inexact, so everything below is the conservative reading: a call resolves only when it names a method of the same class, or a name unique to one class in the project. An upper bound that follows every same-named method agrees on every comparison here.

The total is the same

First result, and it concedes the objection completely. The whole handling path, all calls followed, at the sixtieth rule:

SpringOfficeFloor
Path complexity, growth per rule3.033.25
Difference (Spring minus OfficeFloor)-0.22, interval [-0.57, +0.14], includes zero
Worst single method in the path16.617.6

The two architectures accumulate complexity at statistically identical rates. The additive architecture is not simpler. Its worst function is not smaller. In the blind run it is slightly worse. Anyone who claimed OfficeFloor produces less complexity was over claiming, and that includes the earlier posts in this series.

The unit of change is not the same

Second result. Charge each node only what it reaches, and ask what a developer must hold in their head to change one rule.

Rules implemented515304560
Spring, complexity per change266096138201
OfficeFloor, complexity per change35687

Spring's unit of change grows by about 3 complexity points per rule, in a straight line, for sixty rules. OfficeFloor's grows by 0.09, which against the 3.0 the system as a whole is absorbing is a rounding error. The gap in growth rate is 2.94, interval [2.76, 3.12]. It replicates in a second run under a different test protocol.

The system takes on the same complexity either way. What differs is how much of it you have to face at once.

Where the complexity went

It went into new places to put things. Counting the steps in the create pipeline:

Rules implemented1153060
OfficeFloor, nodes in the path411.115.321.8
Spring, nodes in the path1111

Spring stayed at one node at every checkpoint of every chain, in both runs. It never had anywhere else to put a rule. That is not a criticism of the agent. Nothing in the framework offers a second place, so the handler is the place.

The OfficeFloor nodes are not empty ceremony wrapped around a shared blob. Between 50 and 60 percent of what each node reaches is reachable from that node and from no other. Half to two thirds of each of OfficeFloor step's logic belongs to it alone.

Honest limits

The two arms are measured asymmetrically, and deliberately. OfficeFloor's nodes come from a declared wiring file. Spring has no per-rule node to declare, so its single node is the handler. That asymmetry is the phenomenon, not a thumb on the scale: an architecture earns extra nodes only by actually having separable rules, and Spring never earned one in 1,200 agent sessions.

A fair objection remains. A developer changing one Spring rule does not necessarily read all 201 points of complexity. True. But they cannot know which part matters without reading it, because nothing in the code marks where one rule ends and the next begins. The number is the price of not knowing which part matters.

The lesson

The previous framing was that the additive architecture keeps its handler flat. That framing was measuring the front door, and it would not have survived a determined reader with the wiring file open.

The finding that survives is better. Both architectures absorbed the same complexity at the same rate. After sixty rules, changing one rule in Spring meant facing 201 points of it. In OfficeFloor, 7. Architecture did not reduce the complexity of the system. It decided how much of it you meet at once, and that turns out to be the number that matters.