Sunday, 23 August 2026

The Cold Reader Understood Spring Better

A series on how software architecture shapes AI-driven code degradation. This post reports the strongest result against the position these posts have been arguing. It replicated in both runs, so here it is.

At every phase boundary the harness runs a second agent that has never seen the code. It gets one read-only question:

Read the code as it currently stands and answer without changing anything.
List, in order, every validation and side-effect that is applied when
creating an owner via POST /api/owners.

Its answer is scored against a fixed list of the rules actually implemented. Same question, same list, same checkpoint, both architectures. The expectation, written into the harness before any of this ran, was that the pipeline arm would win easily: its functions are named, wired in order, and enumerable.

It lost.

The cold reader recalled more of Spring

Spring minus OfficeFloor, recall of the rule list, with intervals over chains:

After N rulesBlind runTests-visible run
10+0.051 [0.024, 0.075]+0.058 [0.027, 0.090]
20+0.051 [0.020, 0.085]+0.034 [0.005, 0.064]
30+0.088 [0.015, 0.185]+0.063 [0.007, 0.134]
40+0.046 [0.002, 0.090]+0.010, includes zero
50+0.093 [0.039, 0.144]+0.039, includes zero

Positive means Spring was easier to summarise. Nine of the ten comparisons point the same way and most exclude zero.

And it read twice as much code to do worse

Tokens the cold reader pulled in before answering, averaged over all probes:

RunSpringOfficeFloor
Blind83,000197,000
Tests visible107,000207,000

The difference excludes zero at nearly every probe point in both runs. Reading the distributed version costs about twice as much and yields a less complete answer.

Why

The likely mechanism is unglamorous. Spring's rules are all in one method. Open it and they are in front of you, in order, whether or not you understand them. OfficeFloor's rules are twenty-two wired steps across twenty-two files, and enumerating them means traversing the wiring and opening each one. A reader working to a budget stops early and misses the tail.

Concentration makes a system easy to summarise. That is a real property and it is worth something.

What it does not mean

Summarising is not changing. The same experiment measures what it costs to change a rule rather than describe one, and the answer runs the other way: after sixty rules, changing one rule in Spring means facing about 200 points of complexity, in OfficeFloor about 7. Reading to enumerate rewards putting everything in one place. Changing safely rewards keeping things apart. The two pull in opposite directions and this experiment now has evidence for both.

So the honest summary of the architecture question, after two runs and 2,400 agent sessions, is a trade and not a win. The concentrated design is cheaper to survey and more expensive to modify.

Caveats, because this one deserves them

Recall here is a keyword hit-rate, not a graded assessment. It rewards a reader who mentions the right terms, and terms are exactly what sits densely in one Spring method. A properly graded answer might narrow the gap or close it. The comment in the harness config has said "grade properly offline" since the day the probe was written, and that has not been done.

The question also asks about one endpoint, which is precisely the surface the concentrated arm concentrates into. A question spanning several endpoints might invert the result.

Both caveats were true before we saw which way the number went, which is the only reason it is worth publishing them alongside it.

Tests Are Not How the AI Gets It Right. They Are How You Find Out It Didn't.

A series on how software architecture shapes AI-driven code degradation. This post is about a question the data answers unusually cleanly: what is a regression suite actually buying you when an AI writes the change?

Run sixty accumulating rule changes through a coding agent that cannot see any test except the one for the rule in front of it. Then run the same sixty with every accumulated test visible. The gate scores both the same way, against the full suite, so the measurement never changes. Only the agent's knowledge does.

Finding: regression is specific

Here is the number that surprised us. Across 600 blind checkpoints per architecture, the checkpoints that unintentionally broke an existing rule were:

ArchitectureWhere damage occurred
OfficeFloorrule 39 in 1 chain, rule 51 in 10 of 10 chains
Springrule 39 in 1 chain, rule 51 in 10 of 10 chains

That is the entire list. Fifty-eight of sixty rules produced no unintended damage at all, in either architecture, with no safety net. One rule defeated every single chain in both arms.

The risk is not spread evenly across changes. It is a property of which change. Rule 51 required a new owner's membership level to be capped relative to their household, which meant reaching into behaviour three earlier rules had established. Twenty out of twenty blind chains broke something doing it.

When providing the full tests, the agents successfully passed all tests.  One Spring chain was an exception, breaking four rules at that same rule 51.

What this actually means for testing

The tests earned nothing as a design aid, and nothing as a correctness aid for the change in hand. They earned their keep purely as detection. So the useful question is not "do I need tests" but "what kind of change am I making".

Adding new behaviour was safe to write unverified, indefinitely. Fifty-eight rules, sixty checkpoints deep, no damage.

Revising existing behaviour is where it falls apart. One rule that reached across several established behaviours broke every chain that attempted it.

That reframes the prototype versus maintenance intuition. It is not that young systems are safe and old ones are risky because of age. It is that young systems mostly add, and mature systems increasingly revise, and revision is the operation the agent cannot verify for itself.

One honest limit on the setup. The blind agent was not working with no tests at all: it always received the current rule's own test, which is an unusually precise specification. So this experiment shows that a test-grade description gets the change right, and that the accumulated suite is what stops you silently damaging everything else. It does not show that a vague ticket would have been enough.

We Showed the Agent Every Test. It Wrote the Same Code.

A series on how software architecture shapes AI-driven code degradation. Every previous post ran the agent blind. This one answers the obvious objection to that, and the answer was not the one expected.

The obvious rebuttal to every result in this series: of course the code degrades when you hide the tests. Show the agent what it must not break and it will not break it. Furthermore, it will probably write better code along the way.

So we ran the whole experiment again with the full accumulated suite visible at every checkpoint. Same model, same sixty rules, same two architectures, another 1,200 agent sessions. The agent could see every test it had ever had to satisfy.

Half of that rebuttal is right. The other half is not.

Nothing structural moved

Across roughly thirty-five structural and impact measurements, in both arms, the difference between protocols was statistically indistinguishable from zero. The concentration numbers, which are what this series is about:

Spring, growth per ruleBlindTests visibleDifference
Complexity per change3.0282.953[-0.20, +0.35]
Handler class weight1.7691.621[-0.21, +0.45]
Handler erosion0.003260.00340[-0.0026, +0.0023]

Every interval contains zero. The god handler is created in Spring at the same rate whether or not the agent can see what it is putting at risk.

The agent did not work differently either

This is the part that surprised us. Total effort was flat: turns, cost and output tokens all within three percent between protocols. Deletions actually fell when the tests were visible, significantly so in Spring, 538 lines per chain down to 416. Refactoring language was rare and showed no pattern. The single most aggressive restructuring in the sample, one that replaced three pipeline steps with one and deleted three functions, happened in the blind run.

Most telling: across 2,400 session transcripts, in both protocols and both architectures, the phrases "might break", "could break" and "cannot verify" appear zero times. Not rarely. Never.

The blind agent is not cautious. Tests it cannot see are not uncertainty it weighs and discounts. When the tests are visible, they do not become design pressure either. The agent runs them, fixes what is red, and writes the same shape of code.

What did change

Having a regression suite (all tests) provides a signal for correctness.

OfficeFloor / SpringBlindTests visible
Final-phase checkpoints fully green33% / 27%96% / 93%
Rules broken unintentionally31 and 370 and 4
Chains finishing clean3 of 10, 2 of 1010 of 10, 9 of 10

The price of correctness

Carrying a suite you must read is not free. In Spring, per-checkpoint cost rose by 0.0025 dollars per rule and cache reads by 3,400 tokens per rule as the suite grew (both intervals excluding zero). OfficeFloor showed the same direction without significance. A regression suite is context, and context is billed.

The lesson

A regression suite is detection, not prevention. It will tell you the agent broke something. It will not make the code any better structured, because the agent does not treat it as a signal about design. It treats it as a list to satisfy.

If you are running agents against a mature code base and hoping your test suite keeps the architecture healthy, this experiment says plainly that it will not. It keeps the behaviour correct, which is worth a great deal, and it does nothing whatsoever about the shape of what you are accumulating.

Saturday, 22 August 2026

The Test Suite Crashed. The Harness Called It Regression

A series on how software architecture shapes AI-driven code degradation. This one is not about architecture. It is about a bug in our own measurement, and what it took to find it.

The latest run showed the agent the whole regression suite. The idea was to answer the obvious objection to the blind runs: of course the code degrades when you hide the tests.

The results table said something strange. OfficeFloor, the arm that had broken almost nothing when the tests were hidden, had now broken 143 previously passing rules. Spring, in the same run, broke 34.

That is backwards. More information should not break more rules. And the same table said OfficeFloor passed 594 of its 600 checkpoints outright.

An arm cannot be near perfect and catastrophic at once. One of the two numbers was lying.

What a regression count actually is

At each checkpoint we run the full accumulated suite. We keep the set of tests that passed. At the next checkpoint we run it again and compare. Anything that was passing and is now not passing is a regression.

regressions = prior_passing - now_passing

That is the whole rule. It is set subtraction. It has a failure mode that is easy to miss, and that failure mode is the subject of this post.

Three checkpoints, two chains

All 143 came from three checkpoints, in two of ten chains. Every other OfficeFloor checkpoint in the run was clean.

ChainCheckpointRuleTests runRegressions counted
452identity-key-v2056
816customer-code-city023
853audit-event064

Look at the tests-run column. Zero. Not "some failed". None ran.

Spring had one of these too, in a third chain. It accounted for 30 of Spring's 34. Four crashed checkpoints in the whole run, out of 1,200.

An empty result set means nothing is in now_passing. So everything in prior_passing was reported as a regression. The longer the chain had survived, the bigger the phantom. Checkpoint 53 had 64 accumulated rules, so it invented 64 broken ones.

What actually happened

The build log is unambiguous.

[ERROR] The forked VM terminated without properly saying goodbye.
        VM crash or System.exit called?
[ERROR] Process Exit Code: 134
[ERROR] Crashed tests: ...acceptance.Cp53Tests

Exit 134 is SIGABRT. The JVM that Surefire forks to run the tests died. It produced no reports, because it never got far enough to write any.

Nothing was wrong with the code. The agent's own session at that checkpoint reported 65 tests run, 0 failures, BUILD SUCCESS. The next checkpoint in the same chain passed 66 of 66 with zero regressions, and no repair step in between. The code was fine before the crash and fine after it. Only the measurement died.

Why the harness could not tell

Our gate ran the test command and then parsed the Surefire XML reports. It never looked at the exit code of the test command.

So a crashed run and a clean run with no tests selected produced identical records: build compiled, zero results, no error. The scoring code took the empty map at face value and did arithmetic on it.

This is the general shape of the bug, and it is worth stating plainly. Absence of results is not evidence of failure. A measurement harness that cannot distinguish "I measured nothing" from "I measured zero" will eventually report its most dramatic finding at exactly the moment it knew the least.

The fix, and the part that is easy to get wrong

Classify the run before scoring it. Two detectors, both narrow. No results at all, when the build compiled, cannot be legitimate: at checkpoint K the suite always contains at least checkpoint 1's test. And a Surefire fork-death marker in the console, which is how a partial crash announces the classes that never reported.

Then retry. But only in one direction.

An aborted run is retried, up to three times. A run whose tests merely failed is never retried. That asymmetry is the important line in the whole change. A harness that retries failures until they go green launders exactly the regressions the experiment exists to count. Flakiness is not a reason to run the dice again. It is a reason to know which dice you rolled.

If every attempt aborts, the checkpoint records no verdict at all. Its correctness fields are blank. Missing data, not a score. The analysis drops those rows from every correctness number, keeps their structural metrics, which the crash never touched, and prints the excluded checkpoints by name so the hole is visible in the output rather than absorbed into it.

The second bug, hiding behind the first

Fixing that surfaced a smaller version of the same mistake.

If a checkpoint has no verdict, the comparison set has to carry forward. The next scored checkpoint then measures across the hole: two changes, one diff.

Some checkpoints are mutative. They are required to change a prior rule, so their prior tests are expected to stop passing. Those are intended, and each checkpoint declares which rules it revises.

Three of the four crashed checkpoints were mutative. The exemption was read only from the current checkpoint, so each crashed checkpoint's mandated changes reappeared as unintended breakage on the checkpoint after it. Seven phantom regressions on a checkpoint that was simultaneously reported as fully passing. Passing and broken at once, again, one layer down.

The fourth crashed checkpoint was additive, revised nothing, and left no phantom behind it. That is the control for this bug, sitting inside the same run.

The exemption has to travel with the comparison set. Carry the state, carry its caveats.

What the numbers actually are

As reportedCorrected
OfficeFloor, rules broken1430
Spring, rules broken344
OfficeFloor, clean chains8 of 1010 of 10
Spring, clean chains8 of 109 of 10

One genuine regression survives in the entire run of 1,200 agent sessions. A Spring chain at checkpoint 51 broke four rules, and the same checkpoint failed its own gate. Broken rules and a failing gate, together, in one checkpoint. That is what a real regression looks like in this data, and it is worth noting how different the two phantoms looked. The first reported a whole suite in ruins while its arm passed 594 of 600 checkpoints. The second reported broken rules on a checkpoint that was simultaneously fully green. Both were incoherent before anyone opened a build log.

Two things this does not change. The structural metrics are untouched: erosion, handler complexity, god-class weight and change impact never went through the gate, so the architecture findings stand exactly as published. And the blind run, the one the earlier posts are built on, contains no crashed gates at all. It re-analyses byte for byte identical.

The lesson

The previous post in this series was about a metric that was statistically impeccable and pointed the wrong way. This one is smaller and more embarrassing. The metric was fine. The plumbing lost four measurements out of 1,200, and the arithmetic turned the gap into the loudest result in the table.

Two rules came out of it, and they are not specific to this experiment.

Never let a missing measurement enter arithmetic as a value. Blank is not zero. Zero tests passing is not the same as no tests run, and the difference is the entire finding.

And be suspicious of your own most dramatic number, especially when it is inconvenient for the thing you are arguing. 143 was the single most interesting figure in the run. It was the only one that was not real.

The fix is in the harness, along with the detection that repairs already collected runs without re-running the agent. Every number above is reproducible from the published branches.

Think a Better AI Would Change the Result? The Experiment Is Yours to Run

Every time I publish results from the architecture-degradation experiment, the same objection arrives. A better AI would just refactor the Spring controller each time. Your effect would vanish.

It is a fair question. It is also an empirical one. It is not settled by me arguing in a comment thread. It is not settled by you asserting it either. It is settled by running the experiment. So I have made that easy. The whole harness is public. Swapping the AI model is a single command-line flag.

So do not argue it with me. Run it. Everything you need is here. Run it yourself with a different AI model.

Why the model was fixed on purpose

The experiment holds the coding agent fixed. It makes architecture the independent variable. That is the whole design. Let both the model and the architecture move at once, and you cannot attribute the result to either. Spring versus OfficeFloor was the thing under test.

But fixed for the published run does not mean baked in. Point the harness at whatever AI you think is better. It runs the identical experiment. Same 60 accumulating change checkpoints. Same blind grading. Same isolation. Same metrics. Only the agent changes. That is exactly the variable the better AI model objection is about.

python -m harness.run_experiment \
  --config config.yaml \
  --test-mode blind \
  --model your-better-model \   # the only change vs. the published run
  --run-id your-better-model

The distinction that actually matters

Here is the part most versions of the objection miss. It is the difference between the intercept and the slope.

A better model may well do each change better. That lowers the intercept. But the claim under test is not about any single change. It is about the slope. As change after change lands on the same subsystem, does complexity keep concentrating into one god method and one god class?

The prior benchmark work is a useful clue. Better prompting lowered the intercept. It did not flatten the slope. So the real question is simple. Does raw model capability behave any differently? Or is concentration a property of the architecture, largely independent of how clever the agent holding the pen happens to be?

That is what your run would measure. The doc tells you which numbers to read. The structural slopes. impact_composite. entry_cc. wmc_max. Handler scoped erosion. Each with its confidence interval. And it shows you how to compare them to mine.

Both outcomes are a real result

I am genuinely fine with either way it lands.

  • The effect holds. A stronger model still lets Spring concentrate while OfficeFloor stays flat. That is evidence the effect is architectural. It is not a quirk of one model.
  • The effect weakens. A stronger model refactors the Spring hotspot each time and flattens the slope. That is evidence capability can substitute for architecture. It also answers a good question. How good does the AI have to be before architecture stops mattering?

Both are publishable findings. Neither is something I have to defend in a comment section. That is the point of putting it in a harness.

Why I am handing you the keys

Three reasons. I will be honest about all of them.

  • It removes my bias. I built OfficeFloor. So run it yourself. Use your preferred model. If you get the same shape, that is worth far more than me running it again.
  • It is independent replication, for free. The evolve branches carry raw data only. Anyone can re-derive every number with python -m harness.analyze. Push your branches back. Then the result is checkable by strangers.
  • It does not cost me your tokens. A full run is real money. It is about a week of wall-clock. If you are confident a better model changes the answer, you are the right person to spend that. The doc has a cheap smoke-test-first ladder so you do not find out the hard way.

If you think a smarter AI erases the effect, you might be right. I would like to know. The experiment is sitting there. The instructions are written for exactly this.

Run it yourself with a different AI model.

Bring your own model. Share your branches. Let the data settle it.

The Metrics, Revisited

A companion to the series on how software architecture shapes AI-driven code degradation. This one updates the earlier reference on the metrics. Read that first for the base definitions. This post records what changed.

The first version of this experiment listed thirteen metrics. Running it again, longer and harder, taught us that some of them were measuring the wrong thing, that one of them was actively misleading, and that our headline test was not really a test. So the metric set changed. Here is the update, in one place.

Erosion is now three numbers, not one

The first run had one erosion metric, measured across the whole application. It failed. It could not separate the two architectures, and in the longer run it gave the backwards answer. We took that apart in the leaf-function post. The short version is that whole application erosion is blind to where complexity sits, and it gets swamped by branchy little algorithms both architectures share.

So erosion is now reported at three scopes.

ScopeWhat it coversVerdict
Whole appEvery production functionReported for comparison with the source benchmark. Not decisive. Gave the backwards answer here.
Touched file subsystemOnly files changed since the startNarrower, but still leaf algorithm heavy.
Handler class onlyThe endpoint's own classThe clean concentration signal. Spring's handler erodes. OfficeFloor's does not move.

The lesson generalises. When a metric fails, the fix is often not more data. It is a narrower scope, aimed at the thing you actually asked about.

A new metric: structural impact

If whole application erosion measures the wrong thing, we needed something that measures the right thing. That is the change impact score, built in its own post, on the idea from the cohesion post.

It does not measure how complex the new code is. It measures how much existing, tangled code a change had to disturb, weighted by how tangled that code already was. A new isolated unit costs almost nothing. Surgery on a god method costs a lot. It comes in three parts.

  • Mutation impact. The cost of editing existing functions.
  • Addition impact. The cost of adding new functions, floored so that fragmentation is never free.
  • Composite. The two together. This is the headline number.

Each part is reported three ways. Over all changes. Over additions only. Over mutations only. That split is what let us show that changing an existing rule, not adding a new one, is where the architectures pull furthest apart. That is the subject of the last post in the series.

A better test: difference of slopes

This one is a method fix, and it matters more than it sounds. The first run compared each architecture's degradation slope against zero, one arm at a time. That is not a test of whether the two arms differ. Two error bars that look far apart are not the same as a measured difference.

So now we test the difference directly. We fit a slope to each arm, then bootstrap the gap between them. If that gap's confidence interval excludes zero, the two architectures really do degrade at different rates. This is the headline statistic now, and every decisive claim in the series rests on it.

MetricSpring minus OfficeFloor slope95% CI
Change impact (composite)359[234, 507]
Endpoint handler complexity0.122[0.070, 0.181]
God class weight1.07[0.83, 1.28]
Handler scoped erosion0.0033[0.0016, 0.0049]

All four exclude zero. These are the metrics that carry the result.

The honesty ledger

The same test also told us what to stop claiming. This is the part that matters most for trust, so we state it plainly.

Whole application erosion is demoted. It stays in the results for comparison, but it is no longer a decisive statistic, and here it pointed the wrong way.

Several process metrics did not separate the arms at all. The agent's cost per change did not diverge on the between arm test. Nor did its comprehension cost, the number of packages a change reached into, or temporal coupling. Their difference intervals include zero. So we make no claim that the additive architecture is cheaper to grow or less coupled over time. The data does not support it.

And one nuance we will not bury. The additive architecture disturbs less existing code per change in absolute terms. But that number grows faster for it than for the mutative arm. Its blast radius is smaller, yet not perfectly flat. Additive is not free.

Does the new metric hold up

A metric you invented is only worth something if it predicts something you did not build into it. So we checked the change impact score against signals it never sees. It is computed from the code diff alone. It has no access to what the change cost the agent or whether it broke a test.

It predicts all of them. Within each architecture, a change the metric scores as high impact independently cost the agent more money, took more model time, and forced more re-reading of the code. Changes that broke a rule the agent was never asked to touch scored roughly seven to ten times higher than changes that broke nothing. The score tracks real difficulty and real damage, measured independently of how it is computed.

Reproducibility

Every structural metric here is recomputed from the committed code and its history. That means a new metric can be applied to a past run without ever re-invoking the agent, which is exactly how the change impact score was applied to a run that predated it. The correctness numbers come from a black-box acceptance suite the agent never sees. The harness is on GitHub, and the results in this series all come from a single run.

The full argument, and what it means for building software with an AI agent, is in the hub post.

Adding Features Is Easy. Changing Them Is the Test.

Fifth in a series on how software architecture shapes AI-driven code degradation. The previous post built a score for the cost of a change. This post spends it on the change that matters most.

Most benchmarks measure adding. Give the agent a fresh task. See if it works. Move on.

But adding is the easy part. Real software does not just grow. It changes. A rule you shipped last month gets revised this month. A requirement you thought was settled turns out to be wrong. The bill for software is not paid when you write a feature. It is paid every time you have to change one.

So we measured both. And changing, it turns out, is where the two architectures pull furthest apart.

Two kinds of change

Our experiment feeds each codebase a stream of about sixty rule changes. They come in two kinds.

Most are additive. Add a new rule. A phone format check. A duplicate owner warning. Something that did not exist before.

Some are mutative. Take a rule that already exists and revise it. Change how a telephone number is normalised. Redefine what counts as a duplicate. The rule was there. Now it is different.

That second kind is the real test, and it is the cleanest comparison in the whole experiment. Here is why. When the change is mandated, both arms must make the exact same change. The requirement is identical. Neither arm gets to choose an easier path. So any difference in cost is not about the task. It is purely about the architecture. Same required change. Different bill.

It would be tempting to set these aside. The change was forced, you might say, not the architecture's fault. But that misses the point entirely. The force is equal on both arms. What differs is what each one has to disturb to comply.

The bill

Here is what a typical change costs, scored by the change impact metric from the last post. Lower is better.

Add a new ruleChange an existing rule
OfficeFloor (additive)2382,222
Spring (mutative)3,21411,644

Typical change, measured as the median change impact score across the run.


Read it in two directions.

Down each column, one number stands out. Changing a rule costs far more than adding one. For OfficeFloor, about nine times more. For Spring, about four times more. Changing is the expensive part for both architectures. That confirms the whole premise. The cost of software is in the changes, not the additions.

Across each row, the architecture gap is stark. Spring pays multiples more than OfficeFloor whether it is adding or changing. And one comparison is worth pausing on. It costs OfficeFloor less to change an existing rule than it costs Spring to add a brand new one. The additive architecture's hardest task is cheaper than the mutative architecture's easiest one.

An honest wrinkle

The gap does not widen on changes. It narrows. On additions, Spring pays about thirteen times more than OfficeFloor. On changes, about five times more.

That is not a problem for the thesis. It is the thesis being fair. When a rule must change, OfficeFloor cannot dodge the work either. It has to open up the wired function that owns that rule and edit it. So it loses some of its advantage. Both arms are mutating now.

But look at the absolute numbers, not the ratios. The largest gap in the whole table is on changes. OfficeFloor changing costs about 2,200. Spring changing costs about 11,600. That distance, roughly nine thousand, is bigger than any gap on the additions. Changes are where both the cost and the divergence are highest.

Why the tangle costs

The reason is the same reason the whole series has been circling. Location.

In the additive code base, a rule lives in its own unit. When the rule changes, you change that one unit. The surrounding context is small, so the change impact score stays low.

In the mutative code base, the rule was folded into a growing handler alongside twenty others. All the other rules tangled in beside it, is exactly what the change impact score weights by. You are not just changing a rule. You are disturbing everything it was mixed with.

This is an old idea with a formal name. Architecture researchers measure how far a change can ripple through a system, with metrics like propagation cost and decoupling level. A well decoupled system localises change. A tangled one spreads it. What is new here is watching an AI agent pay that ripple, change by change, in a controlled comparison where the only difference is the architecture.

And it compounds

There is a second effect hiding in the numbers, and it is the worst part for the mutative arm. The handler keeps growing. Every rule that lands in it adds to the surrounding context. So the next revision is weighted by an even heavier tangle. The cost of changing a rule goes up over time, simply because more rules have piled in beside it.

The additive arm has no such spiral. Each rule sits in its own unit, no matter how many other rules exist. Changing one is not made more expensive by the presence of the others. The cost of change stays flat as the system grows. That is the property you actually want when you do not know which requirements will change next. And with an AI agent doing the changing, you rarely do.

The takeaway

Adding features is easy. Any architecture can bolt on new code, and an AI agent will happily do it. The test of an architecture is what happens when the requirements you already built have to change. That is not an edge case. That is the ordinary life of software.

On that test, in this experiment, the additive architecture pays a fraction of the cost, and it does not get more expensive as the system grows. That is the practical shape of the whole result. The full picture is in the hub post.