Saturday, 22 August 2026

The Test Suite Crashed. The Harness Called It Regression

A series on how software architecture shapes AI-driven code degradation. This one is not about architecture. It is about a bug in our own measurement, and what it took to find it.

The latest run showed the agent the whole regression suite. The idea was to answer the obvious objection to the blind runs: of course the code degrades when you hide the tests.

The results table said something strange. OfficeFloor, the arm that had broken almost nothing when the tests were hidden, had now broken 143 previously passing rules. Spring, in the same run, broke 34.

That is backwards. More information should not break more rules. And the same table said OfficeFloor passed 594 of its 600 checkpoints outright.

An arm cannot be near perfect and catastrophic at once. One of the two numbers was lying.

What a regression count actually is

At each checkpoint we run the full accumulated suite. We keep the set of tests that passed. At the next checkpoint we run it again and compare. Anything that was passing and is now not passing is a regression.

regressions = prior_passing - now_passing

That is the whole rule. It is set subtraction. It has a failure mode that is easy to miss, and that failure mode is the subject of this post.

Three checkpoints, two chains

All 143 came from three checkpoints, in two of ten chains. Every other OfficeFloor checkpoint in the run was clean.

ChainCheckpointRuleTests runRegressions counted
452identity-key-v2056
816customer-code-city023
853audit-event064

Look at the tests-run column. Zero. Not "some failed". None ran.

Spring had one of these too, in a third chain. It accounted for 30 of Spring's 34. Four crashed checkpoints in the whole run, out of 1,200.

An empty result set means nothing is in now_passing. So everything in prior_passing was reported as a regression. The longer the chain had survived, the bigger the phantom. Checkpoint 53 had 64 accumulated rules, so it invented 64 broken ones.

What actually happened

The build log is unambiguous.

[ERROR] The forked VM terminated without properly saying goodbye.
        VM crash or System.exit called?
[ERROR] Process Exit Code: 134
[ERROR] Crashed tests: ...acceptance.Cp53Tests

Exit 134 is SIGABRT. The JVM that Surefire forks to run the tests died. It produced no reports, because it never got far enough to write any.

Nothing was wrong with the code. The agent's own session at that checkpoint reported 65 tests run, 0 failures, BUILD SUCCESS. The next checkpoint in the same chain passed 66 of 66 with zero regressions, and no repair step in between. The code was fine before the crash and fine after it. Only the measurement died.

Why the harness could not tell

Our gate ran the test command and then parsed the Surefire XML reports. It never looked at the exit code of the test command.

So a crashed run and a clean run with no tests selected produced identical records: build compiled, zero results, no error. The scoring code took the empty map at face value and did arithmetic on it.

This is the general shape of the bug, and it is worth stating plainly. Absence of results is not evidence of failure. A measurement harness that cannot distinguish "I measured nothing" from "I measured zero" will eventually report its most dramatic finding at exactly the moment it knew the least.

The fix, and the part that is easy to get wrong

Classify the run before scoring it. Two detectors, both narrow. No results at all, when the build compiled, cannot be legitimate: at checkpoint K the suite always contains at least checkpoint 1's test. And a Surefire fork-death marker in the console, which is how a partial crash announces the classes that never reported.

Then retry. But only in one direction.

An aborted run is retried, up to three times. A run whose tests merely failed is never retried. That asymmetry is the important line in the whole change. A harness that retries failures until they go green launders exactly the regressions the experiment exists to count. Flakiness is not a reason to run the dice again. It is a reason to know which dice you rolled.

If every attempt aborts, the checkpoint records no verdict at all. Its correctness fields are blank. Missing data, not a score. The analysis drops those rows from every correctness number, keeps their structural metrics, which the crash never touched, and prints the excluded checkpoints by name so the hole is visible in the output rather than absorbed into it.

The second bug, hiding behind the first

Fixing that surfaced a smaller version of the same mistake.

If a checkpoint has no verdict, the comparison set has to carry forward. The next scored checkpoint then measures across the hole: two changes, one diff.

Some checkpoints are mutative. They are required to change a prior rule, so their prior tests are expected to stop passing. Those are intended, and each checkpoint declares which rules it revises.

Three of the four crashed checkpoints were mutative. The exemption was read only from the current checkpoint, so each crashed checkpoint's mandated changes reappeared as unintended breakage on the checkpoint after it. Seven phantom regressions on a checkpoint that was simultaneously reported as fully passing. Passing and broken at once, again, one layer down.

The fourth crashed checkpoint was additive, revised nothing, and left no phantom behind it. That is the control for this bug, sitting inside the same run.

The exemption has to travel with the comparison set. Carry the state, carry its caveats.

What the numbers actually are

As reportedCorrected
OfficeFloor, rules broken1430
Spring, rules broken344
OfficeFloor, clean chains8 of 1010 of 10
Spring, clean chains8 of 109 of 10

One genuine regression survives in the entire run of 1,200 agent sessions. A Spring chain at checkpoint 51 broke four rules, and the same checkpoint failed its own gate. Broken rules and a failing gate, together, in one checkpoint. That is what a real regression looks like in this data, and it is worth noting how different the two phantoms looked. The first reported a whole suite in ruins while its arm passed 594 of 600 checkpoints. The second reported broken rules on a checkpoint that was simultaneously fully green. Both were incoherent before anyone opened a build log.

Two things this does not change. The structural metrics are untouched: erosion, handler complexity, god-class weight and change impact never went through the gate, so the architecture findings stand exactly as published. And the blind run, the one the earlier posts are built on, contains no crashed gates at all. It re-analyses byte for byte identical.

The lesson

The previous post in this series was about a metric that was statistically impeccable and pointed the wrong way. This one is smaller and more embarrassing. The metric was fine. The plumbing lost four measurements out of 1,200, and the arithmetic turned the gap into the loudest result in the table.

Two rules came out of it, and they are not specific to this experiment.

Never let a missing measurement enter arithmetic as a value. Blank is not zero. Zero tests passing is not the same as no tests run, and the difference is the entire finding.

And be suspicious of your own most dramatic number, especially when it is inconvenient for the thing you are arguing. 143 was the single most interesting figure in the run. It was the only one that was not real.

The fix is in the harness, along with the detection that repairs already collected runs without re-running the agent. Every number above is reproducible from the published branches.

Think a Better AI Would Change the Result? The Experiment Is Yours to Run

Every time I publish results from the architecture-degradation experiment, the same objection arrives. A better AI would just refactor the Spring controller each time. Your effect would vanish.

It is a fair question. It is also an empirical one. It is not settled by me arguing in a comment thread. It is not settled by you asserting it either. It is settled by running the experiment. So I have made that easy. The whole harness is public. Swapping the AI model is a single command-line flag.

So do not argue it with me. Run it. Everything you need is here. Run it yourself with a different AI model.

Why the model was fixed on purpose

The experiment holds the coding agent fixed. It makes architecture the independent variable. That is the whole design. Let both the model and the architecture move at once, and you cannot attribute the result to either. Spring versus OfficeFloor was the thing under test.

But fixed for the published run does not mean baked in. Point the harness at whatever AI you think is better. It runs the identical experiment. Same 60 accumulating change checkpoints. Same blind grading. Same isolation. Same metrics. Only the agent changes. That is exactly the variable the better AI model objection is about.

python -m harness.run_experiment \
  --config config.yaml \
  --test-mode blind \
  --model your-better-model \   # the only change vs. the published run
  --run-id your-better-model

The distinction that actually matters

Here is the part most versions of the objection miss. It is the difference between the intercept and the slope.

A better model may well do each change better. That lowers the intercept. But the claim under test is not about any single change. It is about the slope. As change after change lands on the same subsystem, does complexity keep concentrating into one god method and one god class?

The prior benchmark work is a useful clue. Better prompting lowered the intercept. It did not flatten the slope. So the real question is simple. Does raw model capability behave any differently? Or is concentration a property of the architecture, largely independent of how clever the agent holding the pen happens to be?

That is what your run would measure. The doc tells you which numbers to read. The structural slopes. impact_composite. entry_cc. wmc_max. Handler scoped erosion. Each with its confidence interval. And it shows you how to compare them to mine.

Both outcomes are a real result

I am genuinely fine with either way it lands.

  • The effect holds. A stronger model still lets Spring concentrate while OfficeFloor stays flat. That is evidence the effect is architectural. It is not a quirk of one model.
  • The effect weakens. A stronger model refactors the Spring hotspot each time and flattens the slope. That is evidence capability can substitute for architecture. It also answers a good question. How good does the AI have to be before architecture stops mattering?

Both are publishable findings. Neither is something I have to defend in a comment section. That is the point of putting it in a harness.

Why I am handing you the keys

Three reasons. I will be honest about all of them.

  • It removes my bias. I built OfficeFloor. So run it yourself. Use your preferred model. If you get the same shape, that is worth far more than me running it again.
  • It is independent replication, for free. The evolve branches carry raw data only. Anyone can re-derive every number with python -m harness.analyze. Push your branches back. Then the result is checkable by strangers.
  • It does not cost me your tokens. A full run is real money. It is about a week of wall-clock. If you are confident a better model changes the answer, you are the right person to spend that. The doc has a cheap smoke-test-first ladder so you do not find out the hard way.

If you think a smarter AI erases the effect, you might be right. I would like to know. The experiment is sitting there. The instructions are written for exactly this.

Run it yourself with a different AI model.

Bring your own model. Share your branches. Let the data settle it.

The Metrics, Revisited

A companion to the series on how software architecture shapes AI-driven code degradation. This one updates the earlier reference on the metrics. Read that first for the base definitions. This post records what changed.

The first version of this experiment listed thirteen metrics. Running it again, longer and harder, taught us that some of them were measuring the wrong thing, that one of them was actively misleading, and that our headline test was not really a test. So the metric set changed. Here is the update, in one place.

Erosion is now three numbers, not one

The first run had one erosion metric, measured across the whole application. It failed. It could not separate the two architectures, and in the longer run it gave the backwards answer. We took that apart in the leaf-function post. The short version is that whole application erosion is blind to where complexity sits, and it gets swamped by branchy little algorithms both architectures share.

So erosion is now reported at three scopes.

ScopeWhat it coversVerdict
Whole appEvery production functionReported for comparison with the source benchmark. Not decisive. Gave the backwards answer here.
Touched file subsystemOnly files changed since the startNarrower, but still leaf algorithm heavy.
Handler class onlyThe endpoint's own classThe clean concentration signal. Spring's handler erodes. OfficeFloor's does not move.

The lesson generalises. When a metric fails, the fix is often not more data. It is a narrower scope, aimed at the thing you actually asked about.

A new metric: structural impact

If whole application erosion measures the wrong thing, we needed something that measures the right thing. That is the change impact score, built in its own post, on the idea from the cohesion post.

It does not measure how complex the new code is. It measures how much existing, tangled code a change had to disturb, weighted by how tangled that code already was. A new isolated unit costs almost nothing. Surgery on a god method costs a lot. It comes in three parts.

  • Mutation impact. The cost of editing existing functions.
  • Addition impact. The cost of adding new functions, floored so that fragmentation is never free.
  • Composite. The two together. This is the headline number.

Each part is reported three ways. Over all changes. Over additions only. Over mutations only. That split is what let us show that changing an existing rule, not adding a new one, is where the architectures pull furthest apart. That is the subject of the last post in the series.

A better test: difference of slopes

This one is a method fix, and it matters more than it sounds. The first run compared each architecture's degradation slope against zero, one arm at a time. That is not a test of whether the two arms differ. Two error bars that look far apart are not the same as a measured difference.

So now we test the difference directly. We fit a slope to each arm, then bootstrap the gap between them. If that gap's confidence interval excludes zero, the two architectures really do degrade at different rates. This is the headline statistic now, and every decisive claim in the series rests on it.

MetricSpring minus OfficeFloor slope95% CI
Change impact (composite)359[234, 507]
Endpoint handler complexity0.122[0.070, 0.181]
God class weight1.07[0.83, 1.28]
Handler scoped erosion0.0033[0.0016, 0.0049]

All four exclude zero. These are the metrics that carry the result.

The honesty ledger

The same test also told us what to stop claiming. This is the part that matters most for trust, so we state it plainly.

Whole application erosion is demoted. It stays in the results for comparison, but it is no longer a decisive statistic, and here it pointed the wrong way.

Several process metrics did not separate the arms at all. The agent's cost per change did not diverge on the between arm test. Nor did its comprehension cost, the number of packages a change reached into, or temporal coupling. Their difference intervals include zero. So we make no claim that the additive architecture is cheaper to grow or less coupled over time. The data does not support it.

And one nuance we will not bury. The additive architecture disturbs less existing code per change in absolute terms. But that number grows faster for it than for the mutative arm. Its blast radius is smaller, yet not perfectly flat. Additive is not free.

Does the new metric hold up

A metric you invented is only worth something if it predicts something you did not build into it. So we checked the change impact score against signals it never sees. It is computed from the code diff alone. It has no access to what the change cost the agent or whether it broke a test.

It predicts all of them. Within each architecture, a change the metric scores as high impact independently cost the agent more money, took more model time, and forced more re-reading of the code. Changes that broke a rule the agent was never asked to touch scored roughly seven to ten times higher than changes that broke nothing. The score tracks real difficulty and real damage, measured independently of how it is computed.

Reproducibility

Every structural metric here is recomputed from the committed code and its history. That means a new metric can be applied to a past run without ever re-invoking the agent, which is exactly how the change impact score was applied to a run that predated it. The correctness numbers come from a black-box acceptance suite the agent never sees. The harness is on GitHub, and the results in this series all come from a single run.

The full argument, and what it means for building software with an AI agent, is in the hub post.

Adding Features Is Easy. Changing Them Is the Test.

Fifth in a series on how software architecture shapes AI-driven code degradation. The previous post built a score for the cost of a change. This post spends it on the change that matters most.

Most benchmarks measure adding. Give the agent a fresh task. See if it works. Move on.

But adding is the easy part. Real software does not just grow. It changes. A rule you shipped last month gets revised this month. A requirement you thought was settled turns out to be wrong. The bill for software is not paid when you write a feature. It is paid every time you have to change one.

So we measured both. And changing, it turns out, is where the two architectures pull furthest apart.

Two kinds of change

Our experiment feeds each codebase a stream of about sixty rule changes. They come in two kinds.

Most are additive. Add a new rule. A phone format check. A duplicate owner warning. Something that did not exist before.

Some are mutative. Take a rule that already exists and revise it. Change how a telephone number is normalised. Redefine what counts as a duplicate. The rule was there. Now it is different.

That second kind is the real test, and it is the cleanest comparison in the whole experiment. Here is why. When the change is mandated, both arms must make the exact same change. The requirement is identical. Neither arm gets to choose an easier path. So any difference in cost is not about the task. It is purely about the architecture. Same required change. Different bill.

It would be tempting to set these aside. The change was forced, you might say, not the architecture's fault. But that misses the point entirely. The force is equal on both arms. What differs is what each one has to disturb to comply.

The bill

Here is what a typical change costs, scored by the change impact metric from the last post. Lower is better.

Add a new ruleChange an existing rule
OfficeFloor (additive)2382,222
Spring (mutative)3,21411,644

Typical change, measured as the median change impact score across the run.


Read it in two directions.

Down each column, one number stands out. Changing a rule costs far more than adding one. For OfficeFloor, about nine times more. For Spring, about four times more. Changing is the expensive part for both architectures. That confirms the whole premise. The cost of software is in the changes, not the additions.

Across each row, the architecture gap is stark. Spring pays multiples more than OfficeFloor whether it is adding or changing. And one comparison is worth pausing on. It costs OfficeFloor less to change an existing rule than it costs Spring to add a brand new one. The additive architecture's hardest task is cheaper than the mutative architecture's easiest one.

An honest wrinkle

The gap does not widen on changes. It narrows. On additions, Spring pays about thirteen times more than OfficeFloor. On changes, about five times more.

That is not a problem for the thesis. It is the thesis being fair. When a rule must change, OfficeFloor cannot dodge the work either. It has to open up the wired function that owns that rule and edit it. So it loses some of its advantage. Both arms are mutating now.

But look at the absolute numbers, not the ratios. The largest gap in the whole table is on changes. OfficeFloor changing costs about 2,200. Spring changing costs about 11,600. That distance, roughly nine thousand, is bigger than any gap on the additions. Changes are where both the cost and the divergence are highest.

Why the tangle costs

The reason is the same reason the whole series has been circling. Location.

In the additive code base, a rule lives in its own unit. When the rule changes, you change that one unit. The surrounding context is small, so the change impact score stays low.

In the mutative code base, the rule was folded into a growing handler alongside twenty others. All the other rules tangled in beside it, is exactly what the change impact score weights by. You are not just changing a rule. You are disturbing everything it was mixed with.

This is an old idea with a formal name. Architecture researchers measure how far a change can ripple through a system, with metrics like propagation cost and decoupling level. A well decoupled system localises change. A tangled one spreads it. What is new here is watching an AI agent pay that ripple, change by change, in a controlled comparison where the only difference is the architecture.

And it compounds

There is a second effect hiding in the numbers, and it is the worst part for the mutative arm. The handler keeps growing. Every rule that lands in it adds to the surrounding context. So the next revision is weighted by an even heavier tangle. The cost of changing a rule goes up over time, simply because more rules have piled in beside it.

The additive arm has no such spiral. Each rule sits in its own unit, no matter how many other rules exist. Changing one is not made more expensive by the presence of the others. The cost of change stays flat as the system grows. That is the property you actually want when you do not know which requirements will change next. And with an AI agent doing the changing, you rarely do.

The takeaway

Adding features is easy. Any architecture can bolt on new code, and an AI agent will happily do it. The test of an architecture is what happens when the requirements you already built have to change. That is not an edge case. That is the ordinary life of software.

On that test, in this experiment, the additive architecture pays a fraction of the cost, and it does not get more expensive as the system grows. That is the practical shape of the whole result. The full picture is in the hub post.

Measuring the Blast Radius of Change

Fourth in a series on how software architecture shapes AI-driven code degradation. The previous post argued that we should score a change, not a snapshot. This post builds that score.

We left the last post with a wish list. Score a change, not the code. Weight it by how much existing tangled code the change disturbs. Make isolated new work nearly free. Do not let fragmentation off the hook. And do not be fooled by shallow file splitting.

That is a lot to ask of one number. Here is how it fits into one. Built one decision at a time.

Attempt one: count what you touch

The simplest measure of disturbance is blast radius. How many existing functions did this change modify? Adding a new file touches nothing that was there before. Reaching into ten existing functions touches ten.

This is a real signal, and we track it. But it is too blunt. It treats every function as equal. Modifying a trivial getter counts the same as modifying a two hundred line god method. That is clearly wrong. Not all disturbance is equal.

Attempt two: weight by complexity

So weight each touched function by its complexity. A change that edits a complex function, and edits a lot of it, scores higher than a small poke at a simple one. Multiply complexity by the number of lines changed.

This is better, but it repeats the exact mistake from the erosion post. It charges for the complexity of the changed function itself. A lone soundex algorithm in its own class is complex. Under this scheme, writing it scores high. Even though it is isolated and disturbs nothing. We would be punishing clean, separate code again.

Attempt three: weight by the context, not the unit

Here is the fix, and it is the heart of the metric. Do not weight a change by the complexity of the function it touches. Weight it by the complexity of everything around that function. The other methods in the same class.

Call that the surrounding weight. It is the sum of the complexity of the other methods in the class. It stands for the context you must hold in your head to change this code safely.

Now the numbers behave the way intuition says they should. Edit a method sitting inside a heavy god class, and the surrounding weight is large. The change is expensive. Add a method to a fresh, empty class, and the surrounding weight is zero. The change is nearly free. The soundex algorithm, alone in its own class, costs almost nothing. It disturbs no surrounding context. That is correct. That is the whole idea.

This is the difference between measuring the code and measuring the arrangement. The same edit costs a little in a lean class and a lot in a god class. The architecture that keeps classes lean pays less for every change it makes.

Closing the loopholes

An honest metric has to survive people trying to game it. Two holes needed plugging.

The first is fragmentation. If a brand new class has a surrounding weight of zero, then an agent could win by shattering everything into tiny, empty, cohesionless classes. That is the Modular Mirage from the last post. So we put a floor of one on the surrounding weight. A new isolated unit is no longer completely free. It costs a small, real amount, proportional to its own complexity. Genuine isolation is cheap. Endless fragmentation is not.

The second hole is spread. A single change that reaches into many files is less cohesive than one that stays in a few. So we multiply the whole change by the number of files it touched. Scattering a rule across the codebase costs more than keeping it in one place. This is a deliberate trade. It buys resistance to gaming at the cost of a little separating power, because concentrating in fewer files is something the mutative arm happens to do. We made that trade on purpose and we say so.

One more guard. A rename can look like deleting an old function and adding a free new one. So within a change, if a new function's body closely matches a function that just disappeared, we score it as a modification, not a free addition. Edits cannot hide behind renames.

The formula

Put it together. Each function a change touches has a cost. The whole change sums those costs, then multiplies by the number of files it touched.

cost(function) = max(surrounding_weight, 1) × complexity × max(1, lines_changed)

change_impact = files_changed × Σ cost(function)

summed over every function the change touched, where:

  • surrounding_weight = sum of the complexity of the other methods in the function's class (the context you must hold to change it safely)
  • complexity = the function's own cyclomatic complexity
  • lines_changed = lines the change added or edited in the function (a new function has 0 change)
  • files_changed = number of production files the whole change touched

A new isolated function pays a small floor. A big edit to a complex method inside a heavy class pays a lot. A change smeared across many files pays the file multiplier on top. And here is the part that makes it a degradation signal. As the god method's class grows, the surrounding weight grows too. So the same small edit costs more every time. The mutative arm digs its own hole deeper with each change. The additive arm never starts digging.

What it does to the two arms

We fit this to both architectures across the whole run, then bootstrap the difference between them. The between arm difference excludes zero by a wide margin. Spring pays far more structural cost per change than OfficeFloor, and the gap grows over time.

The size of the gap is worth sitting with. By the end of a run, a single change to the mutative code base disturbs on the order of fifteen times more weighted structure than the same kind of change to the additive one. Early on the two are close. The distance opens with every rule.

But is the number real?

It is easy to invent a metric that tells a nice story. The harder question is whether it means anything. So we checked it against signals it has no access to. The metric is computed purely from the code diff. It never sees how much the change cost the agent, how long it took, or whether it broke a test.

It predicts all three.

Within each architecture, a change the metric scores as high impact independently costs the agent more money, takes more model time, and forces more re-reading of the code. The correlations are moderate and clear, and they hold inside each arm, so this is not just an artifact of one arm being harder overall.

The metric versus an independent signalCorrelation, additive armCorrelation, mutative arm
Money the change cost the agent0.690.54
Re-reading of existing code0.600.49
Model time spent0.680.53

The sharpest test is breakage. Some changes broke a rule the agent was never asked to touch. Those changes scored roughly seven to ten times higher on impact than changes that broke nothing. The metric does not just track effort. It flags the changes most likely to quietly break something.

That is the validation. The score is not a story we like. It is a construct that predicts cost, comprehension, and unintended damage, all measured independently of how it is computed.

Standing on older work

The idea of scoring a change by the risk of what it touches is not new, and we do not claim it. The Delta Maintainability Model scores each commit by the risk of the code it changes. Behavioral code analysis has long ranked hotspots by complexity times change frequency. Others have measured the entropy of how changes scatter across a codebase. All of that is prior art, and all of it is close family.

What is different here is small but pointed. We weight a change by the complexity of the code around it, not the code in it. We make isolated new units nearly free by design. And we use it as the deciding measure in a controlled experiment where the architecture is the only thing that changes. The metric is a tool. The experiment is the point.

There is one question this metric sets up but does not answer. Adding a new rule is one thing. But what happens when an existing rule has to change? That is where the two architectures should differ the most, and it is the subject of the next post.

Complexity Is Not the Metric. Cohesion Is.

Third in a series on how software architecture shapes AI-driven code degradation. Last time the standard erosion metric gave the backwards answer. Here is why, and what to measure instead.

The previous post ended on a problem. Erosion failed because it was blind to location. It measured complexity. And complexity, it turns out, is the wrong thing to measure.

This post is about the right thing.

What complexity actually measures

Cyclomatic complexity counts branches. It tells you how tangled the control flow of a function is. That is a real property. It is also a property of an algorithm, not of an architecture.

Here is a thought experiment. Take one rule. Validate a phone number. Say it carries a complexity of 8. You can drop that logic straight into the existing endpoint handler. Or you can put it in its own small function called ValidatePhone. The complexity is 8 either way. Cyclomatic complexity cannot tell the two apart.

But they are not the same. Not for the humans who maintain the code. Not for the AI agent that changes it next. In one version the phone logic is tangled with twenty other rules. In the other it stands alone. Same complexity. Completely different code.

That is the whole point. Complexity is a property of the code. What we care about is a property of the arrangement.

The word for it is cohesion

Cohesion asks a simple question. Does each unit do one thing, or does it do many things mashed together? A cohesive function has one reason to change. A god method has twenty.

This is what the additive versus mutative story is really about. Not less complexity. Better placed complexity. Small, separate, single-purpose units instead of one growing method that every new rule reaches into.

So the honest fix for our failed metric is not a tweak. It is a change of target. Stop measuring how complex the code is. Start measuring whether responsibilities are kept apart.

Why the textbook cohesion metrics do not help

Cohesion is not a new idea. There is a whole family of metrics for it. LCOM, for Lack of Cohesion of Methods. TCC, for Tight Class Cohesion. So why not just use those?

Because they were built for a different shape of code, and ours breaks their assumptions three ways.

First, they measure cohesion through shared fields. Two methods are cohesive if they touch the same instance variables. But our degradation lives in stateless request handlers. They barely have fields. Their dependencies are injected and shared by everything. So field sharing says nothing. Every handler looks cohesive.

Second, the graph variant, LCOM4, joins methods that call each other. The god method calls all of its little helpers. So they all connect into one component. The class scores as cohesive at the exact moment it becomes a god class.

Third, and worst, none of them look inside a single method. Our main failure mode is one method doing twenty things. That is one method. A metric that measures cohesion across methods cannot see it at all.

The classic tools are not wrong. They are aimed at data classes. We have stateless pipelines and bloating handlers. Different problem.

The trap of fake modularity

There is a tempting shortcut here. If cohesion means small separate units, just reward small separate units. Count the files. Count the functions. More small pieces, better score.

That shortcut is a trap, and it has a name. Recent work on AI-generated code calls it the Modular Mirage. Agents happily split code across many files. But file separation is not cohesion. You can shatter a god method into ten fragments that still make no sense apart. That is not clean. It is a different kind of mess.

So the bar is higher than counting units. A good measure has to reward real isolation. It has to refuse to be fooled by scattering. This matters for us specifically. It would be easy to claim OfficeFloor wins simply because it makes more files. That claim would be worthless unless the new units genuinely absorb change on their own.

The reframe

Put it together and a better idea appears. Do not score a snapshot of the code. Score a change.

When a new rule arrives, ask one thing. How much existing, entangled code did it force the agent to disturb? A cohesive system lets you add a small new unit and walk away. An incohesive one makes you reach into a crowded method that many earlier rules already depend on. And the cost of reaching in should scale with how tangled that method already was.

That is the shift. Not the complexity of what you touched. The complexity of what you had to disturb to touch it. Isolated new work should cost almost nothing. Surgery on a god method should cost a lot. Fragmentation without cohesion should not be free either.

That is a lot to ask of one number. It turns out to fit into one. The next post builds it, and then puts it to the test.

Architecture as the Independent Variable, Run 2: The Signal Was Real

The anchor post for a series on how software architecture shapes AI-driven code degradation. This is a direct sequel to the first run.

The first run ended with two loose ends.

The headline metric, erosion, washed out. It could not separate the two architectures. And correctness did not diverge either. Both arms kept passing their tests. We called the structural damage a latent risk. Real, but not yet biting.

So we ran it again. Three times longer. And this time the signal is not latent. It is loud. But not everywhere, and not in the way the first run expected.

The experiment, in one paragraph

Hold the coding agent fixed. Use the same model for every change. Make architecture the only thing that varies. One arm is Spring, where new rules land as methods on a controller. The other is OfficeFloor, where new rules attach as separate wired functions. Both implement the same feature, one create-owner endpoint, under the same stream of about sixty accumulating rule changes. Ten independent chains per arm. Sixty changes each. Two arms. Twelve hundred agent sessions in total. Then measure how each code base degrades.

Result 1: the structural signal is decisive

First, the statistics, done properly. For each metric we fit a degradation slope to each arm. Then we bootstrap the difference between the arms. If that difference interval excludes zero, the two architectures really do degrade at different rates. This is the actual between arm test.




MetricSpring minus OfficeFloor slope95% CIReading
Endpoint handler complexity0.122[0.070, 0.181]Spring's front door bloats. OfficeFloor's stays flat.
God class weight (WMC)1.07[0.83, 1.28]Spring grows a heavy class. OfficeFloor spreads out.
Handler scoped erosion0.0033[0.0016, 0.0049]Spring's handler class erodes. OfficeFloor's does not move.
Change impact score359[234, 507]Each rule costs Spring far more structural disturbance.

Every one of these excludes zero. The concentration story holds. As rules pile up, Spring pushes complexity into one method and one class. OfficeFloor keeps adding small separate units. That is the thesis, and this time it is measured cleanly, with the right test.

Result 2: erosion was the wrong ruler

The first run's headline metric did not just wash out this time. It gave the backwards answer. Whole application erosion said OfficeFloor was degrading faster. It cleared every statistical bar while pointing the wrong way.

That is not a footnote. It is a lesson in metric validity, and it has its own post. The short version is that erosion is blind to where complexity sits, and it gets swamped by branchy little algorithms that both arms share. We take it apart in Why Erosion Washed Out: The Leaf-Function Trap.

Result 3: a metric that measures the right thing

If erosion cannot see where complexity lands, we need something that can. So we built a change impact score. It does not ask how complex the new code is. It asks how much existing, entangled code the change had to disturb, weighted by how tangled that code already was. A new isolated unit costs almost nothing. Surgery on a god method costs a lot.

That is the metric behind the change impact row above. We describe how it is built, and why it resists being fooled by shallow file splitting, in a dedicated post.

It is not just a number we invented and liked. We checked whether it predicts real pain, using signals it does not have access to. It does. A change the metric scores as high impact independently costs the agent more money, takes more model time, forces more re-reading of the code, and is far more likely to break a rule it was never asked to touch. Changes that actually broke an untouched rule scored roughly seven to ten times higher on impact than changes that did not.

Result 4: what did not diverge

Now the honest part. Not everything separated. Several things did not, and we will not pretend otherwise.

The agent did not spend significantly more per change on one arm than the other over time. The comprehension cost did not diverge. The number of packages a change reached into did not diverge. Temporal coupling did not diverge. Their between arm slope differences all include zero. So we make no claim that OfficeFloor is cheaper to grow or less coupled over time. The data does not support it.

There is one more nuance worth stating plainly. OfficeFloor disturbs less existing code per change in absolute terms. But that number grows faster for OfficeFloor than for Spring. Its blast radius is smaller, yet not perfectly flat. Additive is not free.

What it means

Put the significant and the not-significant together and a clear picture forms.

Architecture decisively governs how change accumulates. In the mutative arm, complexity concentrates, the god method grows, and each new change disturbs more entangled code. In the additive arm, it does not. That is real, measured, and statistically clean.

Architecture does not yet govern whether the code breaks. Correctness fell in both arms as the horizon grew, and it fell together. The gap in genuinely unintended breakage was small. The cost of the mutative arm is structural, not yet functional.

This matches what others are finding. Recent work reports that functional correctness is decoupled from structural quality, and that sheer code volume predicts architectural decay. We do not claim those findings as ours. We confirm them, and we add the controlled comparison they were missing. We hold the agent fixed and change only the architecture. That is what lets us pin the divergence on architecture itself.

The series

This post is the map. Each result above has its own deep dive.

A note on reproducibility. All results come from a single run, identified as blind-202608100006. Every structural number is recomputed from the committed code and its history, so any new metric can be applied to the same past run without re-invoking the agent. The correctness numbers come from a black box acceptance suite the agent never sees.


Update, 23 August 2026. Three of the numbers above are scoped to the endpoint's entry handler: handler complexity, handler class weight, and handler-scoped erosion. A reader pointed out the flaw in that scoping, and they were right. A pipeline architecture can keep its first function pristine by pushing work into the second one, so a front-door measurement flatters it.

We measured the whole path, following every method call from every node. Two things came back. OfficeFloor's entry node is complexity 1.3 while its worst pipeline node reaches about 13, and once helpers are included its create path carries the same total complexity as Spring's controller, growing at the same rate per rule (3.03 versus 3.25, difference including zero). The additive architecture is not simpler, and in the blind run its worst single method is slightly worse.

What survives, and is a stronger result than the one it replaces: the same complexity is packaged into one unit or twenty-two. After sixty rules, changing one rule in Spring means facing about 201 points of complexity; in OfficeFloor, about 7. That gap replicates in a second run under a different test protocol and cannot be produced by relocating work downstream, because the measure follows the work wherever it goes.

The change-impact and blast-radius results above are unaffected: they are computed over the whole diff against whole-application context, not over a scoped path.

Full working in The Same Complexity. One Unit or Twenty-Two.