Saturday, 22 August 2026

Why "Erosion" Washed Out: The Leaf-Function Trap

Second in a series on how software architecture shapes AI-driven code degradation. The first run left a puzzle. This post solves it.

In the first run, one number refused to cooperate. Erosion.

Erosion was meant to be the headline. It should have shown the mutative architecture rotting while the additive one stayed flat. Instead it washed out. Both arms rose. The metric could not tell them apart.

So we ran the experiment again. Longer this time. Sixty changes per chain instead of twenty. And erosion did something worse than wash out. It gave the backwards answer. It said the additive architecture, OfficeFloor, was eroding faster than the mutative one, Spring.


That is the opposite of the whole thesis. It was also, as it turns out, a lesson in how a metric can measure the wrong thing while looking perfectly rigorous.

What erosion measures

Erosion comes from SlopCodeBench. The idea is simple. Give every function a complexity mass. Mass is cyclomatic complexity times the square root of its size. Then take the fraction of all that mass that sits in functions above a complexity of 10.

A high erosion score means most of your complexity lives in a few heavy functions. A low score means it is spread thin. It is a good idea. It is also, for this question, the wrong lens.

The backwards answer

We now test the arms properly. We fit a degradation slope to each arm. Then we bootstrap the difference between them. If the difference confidence interval excludes zero, the two arms really do erode at different rates.

For whole application erosion, the Spring minus OfficeFloor slope difference is negative. The interval excludes zero. In plain terms, OfficeFloor erodes significantly faster. The metric is not confused. It is confidently wrong.

Before blaming the architecture, we checked the obvious suspect. The denominator.

Was it the denominator? No.

Erosion is a ratio. A framework like OfficeFloor routes logic through YAML wiring. That would leave less functions (Java) in the denominator. A smaller denominator inflates the ratio. That would be a measurement artifact, not real erosion.

So we measured the total complexity mass at the start and end of a chain (the denominator).

Spring baseOfficeFloor baseSpring finalOfficeFloor final
Total complexity mass109992720132050

The denominators are close. At the start they are within about sixteen percent. At the end they are almost identical. So the ratio is comparing like with like. The backwards answer is not a denominator trick. It comes from the top of the fraction. It comes from which functions crossed the threshold.

What actually drives erosion

Only a handful of functions in each code base ever cross a complexity of 10. So we listed them. Here is a representative chain.

Spring (mutative)OfficeFloor (additive)
addOwner, CC 23 (the endpoint handler)soundexDigit, CC 19
soundex, CC 19FlagPossibleDuplicate::service, CC 14
soundexCode, CC 19soundex, CC 12
normalizeTelephone, CC 12

Look at what fills these lists. Soundex. Phone number formatting. Duplicate detection. These are branchy little algorithms. A soundex coder is a big character switch. Phone formatting is a pile of conditionals. They are complex in any architecture.

Both arms had to implement the same rules. So both arms grew the same branchy helpers. The erosion score is mostly measuring those helpers. It is barely measuring the endpoint at all.

There is one real difference in that table. Spring's list contains addOwner, the endpoint handler, at complexity 23. That is the god method. That is the thesis made visible. But erosion buries it. It is one line among many.

Why this is the wrong lens

Here is the core problem. Erosion is blind to location. It cannot tell a god method from an isolated algorithm. A function at complexity 19 counts the same whether it is a bloated handler or a tidy, single-purpose soundex coder in its own class.

But location is the entire hypothesis. The claim is not that additive code has less complexity. The claim is that additive code puts complexity in small, separate, single-purpose units. A metric that ignores where complexity sits cannot test a claim about where complexity sits.

Worse, the additive style is penalised for good behaviour. OfficeFloor extracts each concern into its own function. When an extracted algorithm is genuinely branchy, it crosses the threshold on its own. Spring, meanwhile, can bury the same logic inside a larger method and split the counting differently. The threshold is a hard cliff at 10. Cross it and your entire mass counts. Stay under it and none of it does. That makes the score jumpy. On one chain Spring scored low. On the next it scored high. One to five functions decide the whole number.

The fix: scope it to the handler

The repair is small. Stop measuring erosion across the whole app. Measure it only inside the endpoint handler's own class. That is the surface the thesis is about. It excludes the shared soundex and phone helpers, which live in their own classes, in both arms.

Now the metric behaves.



ArmErosion slope95% CI
OfficeFloor (additive)0[0, 0]
Spring (mutative)0.0033[0.0016, 0.0049]

OfficeFloor is flat. Dead flat. Its handler class never erodes, because new rules attach as new wired functions elsewhere. Spring climbs and the interval excludes zero. The two intervals do not overlap. Same raw metric. Same data. Correct answer, once you point it at the right surface.

The lesson

A metric can be significant and still be wrong. Whole application erosion cleared every statistical bar. It had a tight slope. Its between arm difference excluded zero. And it pointed the wrong way, because it was answering a different question than the one we asked.

The failure was not noise. It was validity. The metric was dominated by architecture neutral leaf algorithms, and it was blind to the one thing under test. No amount of extra data would have fixed that.

So the discipline is not only "run the statistics." It is "check what your metric can actually see." Erosion could not see location. Our whole hypothesis was about location.

One note to avoid confusion. Other recent work uses the word "architectural erosion" too. Slater's study means something different by it. There, erosion means violating the layers of a fixed hexagonal design. Here, erosion means complexity concentrating into heavy functions. Same word. Different measurement.

Whole application erosion is not useless. It stays in our results for comparison with SlopCodeBench. It is just not the decisive statistic here, and we no longer treat it as one.

The obvious next question is what to measure instead. If erosion cannot see where complexity lands, what can? That needs a metric built for the job. It also needs a better idea than complexity. That is the next post.

Monday, 10 August 2026

AI graded its own homework

A confound in measuring AI code degradation, and why we deleted every test to fix it.

We have been running a long experiment. Take one feature backlog. It is sixty small, ordered changes to a PetClinic REST service. Have an AI agent implement them one checkpoint at a time. Do it into two different codebases.

One arm is built the conventional Spring way, with controllers and services. The other is built with OfficeFloor and its composed-function architecture. The question is not whether the AI can do it. Both arms stay green almost all the way. The question is how the code decays.

The early runs told a clean story. These were the ones we pushed to GitHub. Spring eroded noticeably worse than OfficeFloor.

It was also partly an artifact of our own measurement setup. At least we now believe so. This post is about the confound we found.

The symptom

Structural erosion here is borrowed from SlopCodeBench. It is the share of a codebase's complexity "mass" that lives in functions above a cyclomatic-complexity threshold. Low is good. Complexity is spread thinly across small functions. High is bad. Complexity is piled into a few fat methods.

In the pushed runs, the final-state numbers looked like this.

Arm (pushed run) Final erosion Create-endpoint handler CC True regressions
Spring 16.9 % addOwner grew to CC 27 4
OfficeFloor 10.5 % entry handler flat at CC 1 8

Two things in that table sit oddly together. Spring eroded more. Its addOwner method ballooned into a 27-branch monster. Yet Spring also regressed less. It had 4 genuine regressions against OfficeFloor's 8. An arm that is quietly accumulating complexity is usually the arm that is quietly breaking things. It is not usually the one breaking fewer.

That mismatch was the thread worth pulling.

The cause is a leftover test that became a reward signal

The two arm repositories were forked from a real application. The Spring arm still carried one pre-existing, native test. It was not part of our harness. It was just a test that shipped with the app. It happened to assert owner-creation behaviour.

Here is the mechanism. It is entirely emergent. Nobody designed it.

  1. A checkpoint changes owner-creation behaviour.
  2. That change makes the old native test fail. The build goes red.
  3. The agent sees the red build. It does the reasonable thing. It updates the test to match the new behaviour.
  4. In doing so it re-encodes the current specification as an executable check. That check then silently guards every later checkpoint against regressing that behaviour.

Repeat this sixty times. The Spring arm has now bootstrapped itself a regression suite. Nobody asked it to. Each checkpoint left behind a slightly better executable spec of what owner-creation should do. Every later checkpoint was quietly held to it.

The OfficeFloor arm had no such leftover test. It got no free regression net.

Our harness did watch for the obvious form of cheating. It checks whether an agent edits the injected acceptance suite we use to grade each checkpoint. That tamper rate was zero across the board.

The confound slipped past for a simple reason. The test it was editing was a legitimate native test. Editing it was correct engineering. It just happened to hand one arm a feedback channel the other arm never had.

Why this explains both anomalies at once

Once you see the self-made regression suite, the odd table resolves cleanly.

Start with the regressions. Spring had 4 and OfficeFloor had 8. Spring had extra protection. Its self-maintained test caught behavioural drift. OfficeFloor was unprotected and shipped that drift.

Now the erosion. Spring reached 16.9 % and OfficeFloor 10.5 %. Spring had a fast red/green signal to hill-climb. The cheapest path to green is to add one more branch to the method already in the failing path. So addOwner accreted conditionals. Its complexity climbed through CC 15, 18, 24, then 27. It climbed in lockstep with the feedback loop. The complexity was the cost of chasing a signal the other arm could not see.

Both fingerprints point to the same hidden feedback loop. Lower regressions is one angle. Higher erosion is the other. So the headline was misleading. "Spring erodes worse than OfficeFloor" was not a like-for-like architectural comparison. It was a comparison between an arm with a private oracle and an arm without one.

The fix is to delete every test except the hidden oracle

The arm base branches are now spring-compare-no-tests and officefloor-compare-no-tests. They carry the application and its test dependencies. They carry no pre-existing test suite at all. The only tests that ever run are the harness's own acceptance checkpoints. Those tests behave in two important ways.

They are copied in per checkpoint. They grade the result in isolation. Then they are reset. They never live in the tree the agent edits.

They are also invisible to the agent in blind mode. It sees neither their contents nor their pass or fail.

Now both arms face identical conditions. Implement the change. Get no test feedback of any kind. Native or injected, there is none. Whatever asymmetry remains has to be the architecture. It cannot be an accident of which fork carried which leftover file.

What the fair comparison actually shows

We re-ran blind with all native tests gone. The picture is more honest. In one respect it is more humbling for a tidy thesis.

Spring's erosion is not reliably worse. In fact it is not reliably anything. Two blind Spring chains ran under identical conditions. They landed on opposite structural styles.

Spring, blind Final erosion addOwner CC / lines Methods in the god class
Chain 0 3.65 % 4 / 48 46 (decomposed)
Chain 1 8.65 % 18 / 92 29 (inlined)

With no signal to hill-climb, the agent falls back on its own prior for good code. That prior is a coin-flip. Sometimes it decomposes into many small methods. Sometimes it inlines into one large one. Spring funnels every rule through a single controller. So that one stylistic choice swings a threshold-based metric by more than 2x. The pushed run's dramatic 16.9 % now looks like one draw of a high-variance process. The self-made test pushed hard toward the inline end and amplified it.

OfficeFloor barely moves. Its erosion sits in a 10 to 14 % band across modes and chains. Its create-endpoint handler stays flat at CC 1 to 2 in every run. The composed-function architecture never offers an inline-a-branch shortcut. A new rule must attach as a new function. So there is no hotspot to bloat.

The interesting property is not just a lower mean. It is lower variance. Concentration does not only raise erosion. It makes erosion unpredictable.

One thing is rock-stable across all of it. Every chain breaks the same way. Both arms, both modes, break the identical pair of household-scoring behaviours at the same late checkpoint. None of them recover. That failure comes from a genuine coupling in the problem itself. A membership-level rule interacts with household-duplicate scoring. It has nothing to do with the test setup. It is the one result the confound never touched. It is the one we trust most.

The lesson for anyone benchmarking AI on code

The tests in a repository are not neutral scenery. They are a reward signal. An agent will climb whatever signal it can see.

Asymmetric test presence invalidates cross-variant comparison. One arm may carry tests another lacks. The same goes for a fork or a configuration. If so, you are no longer measuring the thing you think you are. One leftover test is enough.

An agent editing a legitimate test can silently manufacture an oracle. Watching for tampering with your graded tests is not sufficient. A test the agent is allowed to edit still becomes an executable spec. That spec guards against regression.

Reward signals shape structure, not just correctness. The presence of that test did not only change what passed. It changed how the code was written. It pushed the code toward the fat-method shape that chasing red and green rewards.

So we deleted everything. The only judge left is the one the agent cannot see and cannot touch. 

After cleaning, the early results got a lot noisier. But it is the one that measures architecture instead of accident.


The harness, checkpoints, and analysis are part of the ongoing PetClinic-Evolve comparison. Numbers in this post are from the pushed run 202608081920 and the current blind run. Structural metrics are computed over production Java only. They use the same tools and thresholds for both arms.

Sunday, 9 August 2026

God Methods, Small Functions, and Who Gets to Maintain the Code

Early notes from an experiment that holds the AI fixed and lets the architecture vary. One run each so far. These are my interpretations, not settled results.

I have been running an experiment I call PetClinic-Evolve. The idea is simple. Hold the AI coding agent fixed. Make the software architecture the thing that changes. The agent evolves the same application across about sixty accumulating change requests. One arm is ordinary Spring. Requests route through controller methods. The other arm is OfficeFloor. Behaviour is composed from small wired functions. Then I watch how the code decays as the changes pile up.

This post is early. I have one run of each architecture. That is not enough to prove anything. But the first signal is interesting enough that I want to write down how I am reading it.

The first surprise was how similar they were

For the first fifty checkpoints the two arms moved almost as one. The same features landed. Both stayed green. If you had shown me only the pass counts you could not have told the two architectures apart.

Then both broke at checkpoint fifty one. Both introduced a genuine regression. Both quietly broke an earlier rule while adding a new one. So my first honest conclusion is that neither architecture is magic. Given enough accumulating change, something slips in both.

They broke in different ways

This is where it gets interesting to me. The two arms did not just break. They broke in the shape of their architecture.

Spring concentrates complexity. Over sixty changes its create-owner method grew into a god method. It became long and heavily branched. The class around it grew into a god class made of smaller methods. New rules kept getting stacked into the same place.

OfficeFloor spreads complexity out. New rules arrived as new small functions across new classes. The create function itself barely moved. A handful of functions do carry real complexity, such as a Soundex encoder, but that complexity is inherent to the algorithm.

So when each arm broke, it broke true to type. Spring broke inside the crowded method. OfficeFloor broke across the spread of functions.

Spring healed itself. OfficeFloor did not.

Here is the part I keep turning over. Spring recovered within a single checkpoint. Because everything routes through the same method, the very next change passed back through the broken code and fixed it almost by accident. OfficeFloor never recovered. Its broken functions sat off to the side. Later changes added new functions elsewhere and never came back to them. The regression was stranded, and it stayed broken all the way to the end of the run.

Concentration keeps getting re-touched, so it tends to self-heal, but it bloats. Distribution stays small and readable, but a regression can hide in a corner and persist.

I did not expect that trade. It says the same thing that makes OfficeFloor readable. OfficeFloor's many small isolated functions is also what let a fault sit unnoticed. The same thing makes Spring a mess. Spring cramming everything into one method is what kept dragging the fault back into the light.

The question underneath all of it is who can read the code

OfficeFloor kept its functions within human comprehension. They stayed small and local. You can hold one in your head and reason about it.

Spring did not. A create method that long and that branched is past the point where a person reads it comfortably or tests it by hand.

So both arms end up correct for most of the run. But they are correct in very different ways. OfficeFloor is correct and a human can still follow it. Spring is correct, yet only something that can hold the whole tangled method at once can safely change it. Right now that something is the AI.

That is the thought I cannot let go of. An architecture can pass all its tests and still quietly become code that only a machine can maintain.

Where I think this is heading

There is an important limit in these first runs. The agent only ever saw the single test for the change in front of it. It worked blind, with no memory of the sequence and no sight of the earlier tests. That is deliberate. It is how you measure whether an architecture resists silent breakage. But it also means a broken earlier rule stays broken unless later work happens to touch it.

My belief is that a full regression suite would change the correctness story. If the agent could see every accumulated test while it worked, it would notice the failure and fix it. I think it would keep both architectures accurate. I intend to test exactly that next, by giving the agent the full set of tests already built rather than just the latest one.

And if correctness stops being the difference, then all that is left is maintainability. That is the whole point of the experiment for me. With good tests the AI can probably keep both arms working. But Spring stays working only because an AI can comprehend its god method. OfficeFloor stays working and stays readable by people. One architecture becomes dependent on the AI. The other keeps the door open for a human.

How much to trust this yet

  • This is one run per architecture. It is an early signal, not a finding.
  • The agent was deliberately blind and worked one turn at a time with fresh context. A different setup could shift the picture.
  • The hard numbers, the erosion and complexity trajectories with proper analysis, will come in a later post. This one is interpretation.

Even with those caveats, one sentence captures where my head is. The same choice that keeps OfficeFloor readable by a person is what let a fault hide in it, and the same mess that makes Spring hard for a person is what kept repairing it. If that holds up across more runs, the real question is not which architecture the AI prefers. It is which architecture still lets a human stay in the loop.

See later series on measuring change impact for god classes.

Beyond the Source. When an AI Coding Agent Searches Outside Your Project

I gave a coding agent one task and one directory. It went looking across the whole machine. In doing so it showed me the best thing about how these agents work. It also showed me the most dangerous thing.

I run an experiment called PetClinic-Evolve. It holds the AI coding agent fixed. It makes software architecture the thing that varies. It evolves the same application by the different architectures across roughly sixty accumulating change requests. The goal is to measure how the code degrades over time.

For the measurement to mean anything, each change has to be made blind. The agent is handed the current task and the current source. Nothing else. It must not know it is step 46 of a long sequence.

That blindness was much harder to guarantee than I expected. The agent does not treat the project as the edge of its world.

The moment it reached outside

Early in a run, on the very first checkpoint, the agent needed a Java class. That class is generated from an OpenAPI spec at build time. In its fresh working copy the class did not exist yet. Nothing had been compiled. A person might have run the build. The agent did something more resourceful. It was also more unsettling.

find . -name "OwnerFieldsDto.java"        # not in the project yet
find / -name "OwnerFieldsDto.java"        # so search the ENTIRE machine
# found in /home/.../​.local/share/Trash/.../target/generated-sources/...
grep -n "NotNull|Size|Pattern" "$D/dto/OwnerFieldsDto.java"   # read it from the Trash

The class was not where it should be. So the agent widened its search from the project to the root of the filesystem. It found a copy of the generated code sitting in the desktop Trash. That copy was left over from a previous run. I thought I had cleaned it up. The agent read it to understand the shape of the data model. Then it wrote correct code.

Read that again. It was right. The information it needed really did exist on the machine. The agent found it. The agent used it well. This is not a bug in the model. This is the model being good at its job.

One instinct. Two very different outcomes

Helpful? An agent that looks beyond the files you handed it can fill in gaps you did not know you left. It can find the generated source. It can find an example in a sibling module. It can find a dependency's real implementation. It can find the migration script that reveals a column's true type. It can find the test that pins down the exact error message. Much of what makes these agents feel capable rather than merely obedient comes from this. They infer intent from context you never pointed at. Give one your whole repository and your test suite. It will work out what "done" means. Then it will hit it.

Misleading? The same reach picks up whatever happens to be lying around. Not all of it is relevant. Not all of it is current. Not all of it is even yours. A generated file from a different run. An older version of a class. Another project's conventions bleeding in. A stale schema. The agent does not know the difference between the answer and an answer-shaped thing near the answer. It guesses from what it finds. A confident guess from the wrong source produces confidently wrong code. The resourcefulness has no sense of provenance.

Here is the unifying idea. An AI agent treats the whole machine as context. Not the directory you changed into.

Why this was a five-alarm problem for my experiment

For everyday development, an agent that scavenges the machine for clues is a feature. For a measurement, it is an existential threat. In my setup the machine held the answer key.

The reachable filesystem contained three things. Each was one find / away.

  • A previous run's build output in the Trash. That included its full copy of the authored acceptance tests, Cp01Tests.java through Cp60Tests.java.
  • The harness's own test suite. That is 124 authored test files. They define every checkpoint's requirements.
  • The file checkpoints.yaml. That is the entire future sequence of specifications, in order.

The agent had already shown it would run find / when a file was not where it expected. It never went looking for the tests. But it easily could have. If it had, the experiment would have been quietly and unrecoverably invalid. A degradation study whose subject can read the answer key is not measuring degradation.

I had already removed the tests from the project directory. That was not isolation. It only moved them one directory away. The agent treats that filesystem as fair game.

Keeping the experiment valid

The fix came in layers. It went from cosmetic to structural.

1. Give the agent a history-less, sequence-blind workspace

The agent no longer works in the real git worktree. Each checkpoint it gets a fresh copy of the source only. There is no .git, so it cannot read the commit history and see it is one of a numbered series. There is no prior build output. There is one visible acceptance test. That test is renamed to a neutral AcceptanceTest.java with no checkpoint number and no tag. Nothing in the files or their names hints at a sequence.

2. Make the workspace a flat, ordinary-looking project

That sandbox lives in its own directory. It looks like any project called sandbox. There are no run identifiers in the path. There is no "chain 0 of 3." There is no sibling worktree or .git one cd .. away. There is nothing nearby to stumble into.

3. The real fix. Confine the process to the sandbox

Cosmetics stop an agent from noticing the sequence. They do nothing about find /. The durable fix is to make the rest of the filesystem genuinely unreadable to the agent process. The tests and the spec file cannot be deleted. They are the experiment's source of truth. So the answer is to block access, not to hide the files.

On Linux this is Landlock. It is a kernel feature that lets a process irreversibly restrict its own filesystem access to an allow-list. The agent gets read and write to the sandbox and its toolchain. That means the compiler, the Maven cache, and the CLI runtime. It gets nothing else. Everything outside returns "permission denied."

It is wired to fail closed. Before each turn the harness runs the same confinement over a throwaway shell. It tries to read a few sentinel paths. Those are the test suite and the spec file. If any of them is still reachable, the checkpoint is refused. If Landlock is not available on the host, the checkpoint is refused. It never runs un-blinded. A silent leak becomes a loud stop.

The proof is my favourite part. In a later confined run the agent again ran find / -name "OwnerFieldsDto.java". No one prompted it. This was the exact move that had reached into the Trash before. This time it returned nothing. The escape was attempted and silently defeated. Across forty-five checkpoints, not one file access landed outside the sandbox.

The takeaway cuts both ways

The instinct to look beyond what you were handed is close to the heart of why these agents are useful. They do not stop at the boundary you drew. They go find the context that makes the task solvable. When the context they find is correct, that looks like competence. Sometimes it looks like insight. When the context is stale, foreign, or off-limits, the same behaviour looks like hallucination or cheating. It is the same behaviour.

So the lesson splits by who you are.

  • If you are building with agents, lean into it. Put the good context on the machine. Give it the real tests, worked examples, generated code, the actual dependency source. The agent will use it. It will get more right than if you had fenced it into a tidy little directory. Its reach is a resource.
  • If you are evaluating agents, assume the reach. Your environment is your prompt. A held-out set that merely sits in another folder is not held out. If your result depends on the agent not seeing something, make that thing physically unreadable. Then verify it, every run. Fail closed when you cannot.

I set out to measure how AI-written code decays over time. Before I could measure anything, the agent taught me a lesson. The boundary of a task is not the folder you point it at. It is everything the process can reach. Draw that boundary deliberately. Otherwise the agent will draw it for you.


Friday, 7 August 2026

Mutation of existing logic showing to cause erosion

The redesigned PetClinic-Evolve experiment has produced its first complete run of each architecture. It is one chain per arm, so it is a first look and not the final verdict. But it already moves the measure that stayed flat for the whole first experiment, and it moves it in the direction the thesis predicts.

The setup, in one paragraph

Hold the coding agent fixed. Vary the architecture. One agent model builds the same app twice. In the Spring version, request logic lands in controller methods. In the OfficeFloor version, each rule is a small function wired together by YAML. Then sixty change requests land on the same endpoint, create owner, and we watch how each code base ages. Every fourth change is mutative: it revises earlier rules rather than only adding. The idea under test is that Spring concentrates the accumulating logic into one growing method, while OfficeFloor spreads it across many small functions and keeps the entry point flat.

The entry handler: the decisive measure

The cleanest number is the complexity of the one function the create endpoint routes through. In Spring that is the controller method. In OfficeFloor that is the pipeline's create function.

CheckpointSpring create handler (CC)OfficeFloor create function (CC)
100
16112
32132
48219
60279

Cyclomatic complexity counts the independent paths through a function. A value around 27 is a method that is genuinely hard to hold in your head. Spring's create handler climbs to 27 and is still rising at the end. OfficeFloor's create function sat at 2 through the first half and ends at 9. Fitted as a trend, the Spring handler grows about 2.7 times faster per change. This is the mechanism in one line. Rules pile into the Spring handler. They attach beside the OfficeFloor one.

Where the complexity lives

Here is the part that a single summary number hides. Both code bases end with three or four functions above the usual complexity threshold. So a blunt erosion ratio looks similar for the two. But the functions that carry the complexity are not the same kind of thing.

Spring's busiest functions at the end:

ComplexityFunction
27the create controller method itself
19a soundex name-coding routine
12a region-code helper

OfficeFloor's busiest functions at the end:

ComplexityFunction
19a soundex digit routine
13the possible-duplicate rule
13a soundex helper
11the telephone formatting rule

In Spring the single busiest function is the front door itself. In OfficeFloor the busiest functions are isolated, single-purpose units, and the front door stays flat. Notice the soundex routine sits at complexity 19 in both arms. That is the inherent complexity of the algorithm, not erosion, and it shows up in both. The difference is that Spring carries a complexity-27 god method on top of that shared cost. OfficeFloor does not.

Erosion, over the whole run

The erosion measure is the share of complexity that lives in functions above the threshold. It stayed at zero for the entire first experiment. In this run it moves.

CheckpointSpring erosionOfficeFloor erosion
80.000.00
160.240.00
320.280.11
480.290.22
600.330.23

Spring erodes early, from checkpoint 16, and settles around 0.33. OfficeFloor stays at zero until checkpoint 32, then rises to 0.23. So OfficeFloor is not immune. The deep mutations in the back half do push it up. But it erodes later, it erodes lower, and it erodes in a spread out way rather than concentrating in the handler.

Correctness and cost

Both arms passed the same number of checkpoints cleanly. That number is dominated by a few of my own tests that were too brittle, which fired the same way in both arms, so I do not read much into the correctness magnitude from this run. I have since made those tests stricter, and the next run will give a correctness picture worth trusting. The early hint is that Spring broke a wider set of earlier rules.

Cost was close. The Spring chain cost about 79 US dollars in agent time. The OfficeFloor chain cost about 85. OfficeFloor did a little more total work, with more functions and more lines, because it spreads the same behaviour across more units. So it is not that one arm did less. It is that the two arms distributed the same job differently.

Honest limits

  • This is one chain per arm. Any single chain can go its own way. The real result needs many independent chains per arm and the confidence intervals across them. That run is underway.
  • The correctness comparison is muddied by brittle tests in this first chain. The structural comparison does not depend on those tests, so it stands.
  • The erosion ratio alone is a poor summary. The entry-handler complexity and the identity of the busiest function are what separate concentration from distribution.

Where this goes

On the measure that could not move in the first experiment, the two architectures now separate clearly, and in the predicted direction. Spring's create handler became a complexity-27 god method that is still growing. OfficeFloor's create function stayed flat at 9, with the complexity pushed out to bounded, single-purpose functions. The structural half of the thesis is looking well supported. The safety half, whether the composed design also regresses less, awaits the run on the stricter tests. When the full multi-chain run completes, the slopes and their confidence intervals will turn this first look into an answer.

A later series looks at better measuring mutation costs.

Improvements to experiment: more checkpoints, mutative checkpoints, blind regression measurements

This is a between experiments post. The first PetClinic-Evolve run gave a clear answer on some measures and a flat non answer on others. The flat parts turned out to be the interesting ones. Here is why they came out flat, and how the next experiment is built to force the question.

The question

PetClinic-Evolve keeps the coding agent fixed and makes the architecture the thing we vary.  A Spring version where request logic lands in controller methods. An OfficeFloor version where each rule is a small function wired together by YAML. Then a long stream of change requests lands on the same endpoint, create owner, and we watch how each code base ages.

What the first run showed, and where it went quiet

The first run walked 20 change checkpoints, with ten independent chains per architecture. It cost about 510 US dollars in agent time. Some signals came through cleanly. The blast radius measures separated between the two arms. So did the growth of the single entry handler, and a measure of how often new changes reopened old code. Those trends pointed the right way.

Two headline measures stayed silent, and that is what prompted the redesign.

  • Erosion washed out. The erosion measure stayed at zero for the whole run, in both arms. No single method ever grew complex enough to trip the measure. The agent kept splitting logic into many small methods, so no one method ever spiked.
  • Regressions were zero. Neither arm broke a previous behaviour. The safety difference the experiment exists to measure never showed up. It was hidden, not absent.

Why those two came out flat

Neither flat result was reassuring. Each one traced back to a choice in the test harness, not to the code being healthy. There were four causes.

1. The run was not long or deep enough

Twenty additive checkpoints did not push Spring past the point where a large method forms. Complexity did build up, but it stayed spread across many small methods. A measure that waits for one method to grow complex has nothing to report until the pressure is far higher.

2. Every checkpoint only added

Adding is the easy case, and it quietly favours the OfficeFloor design. Adding a brand new rule as a brand new function is exactly what that architecture is good at. Real maintenance is not only adding. It makes changes to previous requirements (i.e. mutating the existing logic of the application). An experiment made entirely of additions never tests the case that hurts most.

3. Regression was almost impossible by design

This was the important one. At each checkpoint the agent could see every previous test. So it had a full checklist of what not to break. With that checklist in front of it, of course it did not break anything.

4. The tests were too soft, and isolation was not tight enough

Some tests only checked that a field was present, not that it held the right value. A presence check cannot notice a wrong value, so it cannot notice a regression. Separately, the coding tool has a memory feature that can write notes between runs, which risks carrying knowledge across checkpoints that are meant to be independent.

How the next experiment is formed

The redesign tackles each cause directly. The thing we vary, the architecture, is unchanged. The instrument around it is rebuilt.

First run limitationRedesign
Too short and shallow, so erosion never had a chance to appear.Sixty checkpoints, three times longer, so pressure on the single handler builds well past the first run.
Only additions, which favoured the addition friendly arm.Mutative checkpoints. Roughly every fourth change now revises earlier rules rather than only adding. Their reach grows from two earlier rules up to six, with deliberate deep changes near the middle and at the end.
The agent saw all past tests, so regression was near zero.Blind regression measurement. The agent sees only the current checkpoint's test. The full set of past tests is used afterwards to check for regression.
Soft tests that only check for presence.Exact tests. Every test asserts a precise value. Computed values such as hashes and check digits are recomputed inside the test. Look ups use small fixed tables shared by both arms.
Possible memory carried between checkpoints.Isolation per turn. Each agent run starts with a fresh, login only setup, so the tool cannot carry notes between checkpoints or between arms.

The mutation, and the rule it needs

The mutative checkpoints are the heart of the redesign. When a checkpoint changes an earlier rule, it provides updated previous tests for the mutation.

The previous checkpoints being mutated are flagged by the checkpoint. A break in a previous checkpoint rule that was not on the list is a clear regression.  This now allows for a safety signal regarding the changes.

An early look, offered with caution

One Spring chain of the new design has run from start to finish. It is a single chain, one arm, and it is not the comparison. But it already shows the instrument now moves where it used to sit still. The erosion measure, which stayed at zero for the whole first experiment, now lifts as the change stream deepens. It rises through the first third of the run and peaks near the first deep mutation.

CheckpointErosion measure
10.00
80.00
160.19
240.22
320.28
400.22
480.19
560.20
600.20

Over the same chain, the single busiest method grew from a complexity of one to nineteen, and the total code grew about thirteenfold. This is one chain, and it is Spring only. It validates the instrument, not the thesis.

Regressions appear now too, and they begin at the first mutations and build up, which is the shape the thesis predicts.

Where this goes

The first experiment was not a failure. It was a calibration. It showed which signals the design could already separate, and it showed exactly which measures needed a harder test before they could speak. The redesign is that harder test. Longer, with real mutation, with regressions made visible rather than assumed away, and with tests strict enough to trust.

The next experiment run is underway.

Tuesday, 4 August 2026

Measuring AI-driven code degradation

Methods companion · Measurement reference Spring OfficeFloor

Measuring AI-driven code degradation

A companion to the results paper. Every metric in the PetClinic-Evolve harness, defined precisely. What it computes, where the number comes from, and what it reads high on.

1

Why the measurement comes first

A degradation study is only as good as its metrics. If the numbers are vague, the conclusion is vague. This companion defines each metric the harness records, so a reader can check the results against exactly what was measured. The results paper reports what the numbers did. This paper says what the numbers are.

Two design choices shape everything below. Both exist to make a per-change measurement trustworthy.

The agent delta is isolated. Each checkpoint produces two commits. The first is the agent commit. Its diff against its parent is exactly what the AI changed, and nothing else. The second is a reset commit that re-normalizes the tree for the next step. Any metric described as "of a change" is computed on that agent diff. So blast radius, coupling, and churn measure the agent's edit, not harness bookkeeping.

Capture now, derive later. A run stores only raw data on the branches. The agent envelope. The raw pass or fail of every test. The committed source at each step. Every metric below is recomputed offline from that history. So a metric can be defined or corrected after a run and re-applied to it, with no new agent cost. Every number was produced this way.

2

Ground rules for the structural metrics

Four conventions apply to every structural metric in Section 3.

  • Production Java only. Structural metrics ignore test code, build files, and configuration. Java under the main source tree is counted. OfficeFloor's YAML wiring is counted separately and never mixed into a Java denominator.
  • Per-function CC and SLOC. A static analyzer (lizard) reports two numbers per function. CC is cyclomatic complexity, the count of independent paths through the function. SLOC is source lines of code. These two feed most of what follows.
  • The complexity threshold is 10. A function is "complex" when its CC exceeds 10. This is the standard Radon threshold, applied identically to both arms.
  • The dynamic subsystem. Some metrics run over the whole application. Others run over the "touched subsystem": every production function in a file changed since the chain's base commit. Scoping this way follows the evolving footprint, and it automatically includes any new class the agent creates. So a growing hotspot cannot hide in a fixed file list.
3

Structural metrics

Erosion ratio · whole-app and scoped

mass(f) = CC(f) × √SLOC(f)
Erosion = ΣCC(f) > 10 mass(f)  ÷  Σall f mass(f)
Measures
How concentrated complexity is. The share of total complexity "mass" that sits in functions above the CC threshold.
Source
lizard CC and SLOC over production Java. Reported twice: over the whole app, and over the touched subsystem (erosion_scoped).
Reads high
When a few heavy functions hold most of the complexity.
In this run
wash Both arms near identical. A ratio normalizes away the concentration it targets. See Section 6.

Verbosity ratio

Verbosity = | CloneLines ∪ PatternLines |  ÷  LOC
Measures
Duplicated and boilerplate-flagged lines as a share of code.
Source
Clone lines from a copy-paste detector (jscpd, strict mode). Pattern lines from a structural search (ast-grep) run against a rule set of wasteful Java idioms. The two line sets are unioned, then divided by lines of code.
Reads high
When code is repeated or padded with boilerplate. Note the index can exceed 1.0, because a clone counts lines on both sides of the pair. It is an index, not a fraction.
In this run
separates OfficeFloor higher at every phase, from its many similar small functions. Neither arm's verbosity grows.

Blast radius per-change family

existing_fns_modified = # already-present functions whose body the diff touched
files_created = # new production files  ·  churn = production lines added / removed
Measures
How much pre-existing code a single change disturbs, and whether it adds or mutates.
Source
The agent diff. Hunk headers (git diff -U0) give the changed line ranges. lizard gives each function's line range. A function counts as modified when a changed range overlaps its body, and its file existed before the change. New files are counted as created.
Reads high
existing_fns_modified is high when a change reaches into working code. files_created is high when a change adds new units instead.
In this run
separates Spring's existing_fns_modified slope excludes zero; OfficeFloor's is flat and it creates ~19 files per chain.

Hotspot CC worst function

hotspot_cc = maxf in touched subsystem CC(f)
Measures
The single most complex production function in the evolving footprint, reported with its name.
Source
lizard over the touched subsystem. The function with the highest CC, ties broken by SLOC.
Reads high
When one function is becoming a god-method.
In this run
overlap Both arms rise; Spring higher, but the confidence intervals overlap.

Weighted Methods per Class (WMC) worst class

WMC(class) = Σmethods m in class CC(m)
wmc_max = maxclass WMC(class)
Measures
The god-class indicator. The heaviest class by total method complexity.
Source
lizard functions grouped by file (one top-level class per Java file), summed per class, maximized over classes in the touched subsystem.
Reads high
When one class carries many rules. This catches what erosion misses: a class that stays tidy method by method while accumulating many methods.
In this run
overlap Both rise; Spring's final WMC is ~2× OfficeFloor's, but the slopes overlap.

Entry-handler CC the front door

entry_cc = CC( f ) where f matches the arm's create-entry pattern
Measures
The complexity of the one function the create endpoint routes through.
Source
An arm-specific regular expression over "file::function". Spring matches OwnerRestControllerV1::addOwner. OfficeFloor matches its designated create-entry, BuildOwner::service. lizard gives the CC.
Reads high
When the single handler absorbs each new rule rather than delegating it.
In this run
separates cleanly Spring +0.25 per checkpoint, OfficeFloor +0.06. The confidence intervals do not overlap. Final CC 6.7 vs 1.9.

Change spread reach

packages_touched = # distinct package directories the diff's production Java touches
Measures
How far across the package tree a single change reaches. A sibling to blast radius, measured by directories rather than functions.
Source
The agent diff. Distinct parent directories of the changed production Java files.
Reads high
When a change is scattered across many packages.
In this run
both near zero Neither arm's spread trends.

Re-edit rate temporal coupling

for each function the change edits, blame its whole body at this commit
re-edit rate = ( body lines authored by an earlier checkpoint ) ÷ ( total edited-function body lines )
Measures
When a change edits existing code, how much of that code earlier rules wrote.
Source
For every production function the change modifies, git blame attributes each body line to the checkpoint that last wrote it. A line counts as "prior" when its author is a checkpoint after the base but before this one. Whole function bodies are counted on purpose. A one-line insert into a large shared method still signals coupling, and line-of-diff blame would miss it.
Reads high
When new rules keep reopening functions that earlier rules grew.
In this run
separates cleanly Spring's slope is positive, OfficeFloor's is negative. The intervals do not overlap. The sign of the coupling trend flips between the architectures.

Function-package stats OfficeFloor-specific

fn_count, fn_nloc_avg, fn_nloc_max, fn_cc_max over the wired-function package
Measures
The size distribution of OfficeFloor's composed-function package.
Source
lizard over a configured package glob (the rest/function tree).
Reads high
The healthy-growth signal is a rising count while the average and maximum function size stay flat. That is addition without bloat.
In this run
Count rises steadily; per-function size stays flat. Consistent with add-not-mutate.

Lines of code size, reported separately

java_loc = production Java SLOC  ·  yaml_loc = OfficeFloor wiring lines
Measures
Raw size. Kept as context, and as the denominator for verbosity.
Source
lizard for Java. A non-blank non-comment line count for YAML.
Note
YAML is never folded into a Java ratio. It is reported on its own so OfficeFloor's habit of spreading logic into wiring cannot distort a Java metric.
4

Correctness metrics

Correctness is scored from the raw pass or fail of every acceptance test, captured at run time. The tests are black-box. They hit the REST API, so both arms are judged by identical externals.

The test taxonomy. There is one test class per checkpoint, named CpNN, tagged so the harness runs checkpoints one through K at checkpoint K. The method-name prefix encodes a category: core, error, or functionality. A test whose own checkpoint is earlier than K counts as a Regression test at K, whatever its category.

Strict / ISO / Core pass correctness gates

strict = all selected tests pass  ·  iso = this checkpoint's own tests pass  ·  core = the core tests pass
Measures
Three tiers of "did it work". Strict is everything. ISO ignores the regression tests and asks only whether this checkpoint's new behavior works. Core asks only the essential path.
Reads high
All three are 1.0 when the checkpoint fully satisfies its suite.
In this run
Strict pass was 1.0 in every phase, both arms.

Normalized Change SWE-CI · range -1 to 1

if passed ≥ baseline:  (passed − baseline) ÷ (target − baseline)
else:  (passed − baseline) ÷ baseline
Measures
Net progress in passing tests from one checkpoint to the next, on a signed scale.
Source
The passing-test set before and after the checkpoint. "baseline" is the prior passing count, "target" the total selected.
Reads
Positive for improvement, negative for regression. The definition is asymmetric on purpose. It punishes a regression harder than it rewards an equal-sized improvement.

Regressions count

regressions = | passing_before − passing_after |
Measures
How many tests that passed before the checkpoint fail after it.
Reads high
When a change breaks previously working behavior. This is the direct safety signal.
In this run
no difference Zero regressions in either arm across 200 checkpoints each.

EvoScore SWE-CI · discounted success

per chain:  ( Σi γi · si ) ÷ ( Σi γi ),  then averaged over chains
Measures
Success across a whole chain, weighted by checkpoint. si is the strict-pass indicator at checkpoint i.
Source
The strict-pass sequence per chain. The discount γ is a parameter; γ ≥ 1 rewards staying green late in the run, when the codebase is largest.
In this run
1.0 at every γ for both arms.

Zero-Regression Rate across chains

ZRR = ( # chains with zero total regressions ) ÷ ( # chains )
Measures
The share of full runs that never broke anything.
Reads high
1.0 means every chain stayed regression-free end to end.
In this run
no difference 1.000 for both arms.
5

Process metrics

These come straight from the agent session envelope, captured at run time. They cannot be recomputed later, so they are stored raw.

Cost, tokens, and duration the agent envelope

cost_usd
Dollar cost of the checkpoint's agent session.
tokens
Input, output, cache-read, and cache-creation token counts. Cache-read is a proxy for how much prior context the model re-read.
num_turns
Tool-use turns the agent took to finish.
duration_ms
Wall-clock time, including tool runs and any waits.
duration_api_ms
Model inference time only. The cleaner "thinking cost" signal.
attempts
Every attempt is recorded, including failed or rate-limited ones and their wait time, so true cost and wall-clock are recoverable.
In this run
Cost and API time fell slightly over each run in both arms. The "comprehension gets more expensive" idea did not appear at this scale.
6

The statistic: degradation slope

A single metric at a single checkpoint is noise. The signal is the trend across the run. The harness reduces each metric to one number per arm, with an interval.

Degradation slope m the headline number

m = OLS slope of ( metric value vs checkpoint index )
reported on the mean curve, with a 95% bootstrap CI resampled over chains
Mean curve
At each checkpoint, average the metric across the ten chains. Fit a straight line to that averaged curve. Its slope is m.
Confidence interval
Resample the ten chains with replacement, 2,000 times. Refit m each time. The 2.5 and 97.5 percentiles are the 95% interval. The interval reflects chain-to-chain variability.
Phase means
Checkpoints one to twenty are also binned into five phases, Start to Final, to show the trajectory in a small table.
How to read a slope A metric separates the two architectures when one arm's interval excludes zero and the other's brackets it. It separates strongly when the two arms' intervals do not overlap. A rising slope with an interval that stays clear of zero is the degradation signal. An interval that straddles zero is a flat metric.

The one caveat that matters. Not every metric can see every difference. Erosion is a whole-application ratio. When one arm adds many small functions, it inflates the denominator at the same time the other arm concentrates complexity in the numerator. Both effects move the ratio the same way, so the ratio cancels the difference. The lesson generalizes. For architectural degradation under AI-driven change, lead with un-normalized, per-change statistics such as blast radius, entry-handler complexity, and re-edit rate. Treat aggregate ratios as secondary. The headline number should be one that can discriminate.

7

Controls that keep a number honest

Three controls protect the metrics from being gamed or contaminated. They are part of the measurement, not a footnote to it.

  • Experimenter-owned tests, reset before scoring. The acceptance tests are black-box and reset to their authored version before the correctness gate runs. An agent that weakens a test cannot produce a false pass. Any edit it makes is recorded, then reverted.
  • A pinned project guide. The leveling document is restored to its base version after every checkpoint, so it can never become accumulating memory across checkpoints.
  • Provenance and a config snapshot. Each run records the model, the tool versions, the agent environment, and a snapshot of the exact metric configuration used. So a metric recomputed later uses the same definitions the run was scored under.

In this run the agent never edited a pinned document and never tampered with a test, across all 400 sessions. The controls were never triggered. Their presence still removes two ways the numbers could have been wrong.


R

References and notes

  1. SlopCodeBench · arXiv:2603.24755. Source of the Erosion and Verbosity definitions, the degradation-slope statistic, and the prompt-intervention arms.
  2. SWE-CI · arXiv:2603.03823. Source of Normalized Change, EvoScore, and Zero-Regression Rate.
  3. Weighted Methods per Class follows the Chidamber and Kemerer object-oriented metric suite.
  4. OfficeFloor · officefloor.net. The graph-of-functions framework used in the OfficeFloor arm.

Companion. The results, with confidence intervals and the full slope table, are in "Architecture as the independent variable". This paper defines the instruments. That paper reports the readings.

PetClinic-Evolve · metrics reference · claude-opus-4-8 · every definition matches the harness implementation and is recomputed from the committed run data.