Wednesday, 23 September 2026

The structural-impact score

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Structural impact: the composite score

impact_composite  ·  ↓ lower is better  ·  defined in this harness (harness/metrics.py)

Structural impact: the composite score

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

One number for how much a change cost the structure. It is blast radius weighted by the complexity of the context that was disturbed. Editing a method inside a heavy god class costs far more here than the same edit inside an isolated unit, which is the whole design.

How it is calculated
cost(f) = max(WMC_other(f), 1) · CC(f) · max(1, Δlines(f))
impact_composite = files_changed · Σf ∈ changed cost(f)

Where. cost(f) is the per function term. WMC_other(f) is the summed cyclomatic complexity of the other methods in f's class, which stands for the context you must hold in your head to change f safely. A brand-new class has WMC_other = 0, which the max(…, 1) floors to 1, so a new isolated unit still costs something. That floor is what closes the fragmentation loophole. Δlines(f) is the changed-line count in f, and for a new function it is that function's own size. files_changed is the number of distinct production Java files the commit touched, applied as a multiplier, so scattering one rule across many classes is not free. The sum runs over every changed function, both modified and new. A within-commit rename is charged as a mutation rather than a free addition when the two bodies have a line-set Jaccard similarity of at least 0.6.

In this harness. metrics.impact_stats parses each touched file at both the previous and the current commit with lizard, matches functions by name, and falls back to body similarity for renames. Note what it cannot see: only lines inside a parsed function body count, so logic expressed declaratively, in a MapStruct expression, in openapi.yml, in schema.sql or in OfficeFloor's wiring, scores zero. Both arms have that escape hatch, so it is not an arm bias, but read a zero as “the logic went where this metric cannot look” rather than as “the change was cheap”.

How to read it. The scale is large and heavily skewed, because it is a product of four terms. The shape of the line matters far more than its value.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact: disturbing existing functions

impact_mutation  ·  ↓ lower is better  ·  this harness

Structural impact: disturbing existing functions

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The half of the impact score that comes from modifying code that already existed. This is the component that carries the discrimination between the two architectures, because the context weight means a mandated rule revision is genuine architectural signal rather than spurious re-touching.

How it is calculated
impact_mutation = files_changed · Σf ∈ modified ∪ renamed cost(f)

Where. cost(f) is the per function term. WMC_other(f) is the summed cyclomatic complexity of the other methods in f's class, which stands for the context you must hold in your head to change f safely. A brand-new class has WMC_other = 0, which the max(…, 1) floors to 1, so a new isolated unit still costs something. That floor is what closes the fragmentation loophole. Δlines(f) is the changed-line count in f, and for a new function it is that function's own size. files_changed is the number of distinct production Java files the commit touched, applied as a multiplier, so scattering one rule across many classes is not free. The sum is restricted to functions that existed at the previous commit and were modified, plus those detected as renames by the 0.6 Jaccard rule.

In this harness. Same single pass as the composite.

How to read it. This is where the two architectures separate most sharply. If the composite moves and this does not, the movement was all in additions.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact: adding new functions

impact_godclass  ·  ↓ lower is better  ·  this harness

Structural impact: adding new functions

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The half of the impact score that comes from new code: new files and new methods added to existing classes. It is reported so the composite's behaviour can be attributed to the right half.

How it is calculated
impact_godclass = files_changed · Σf ∈ new cost(f)

Where. cost(f) is the per function term. WMC_other(f) is the summed cyclomatic complexity of the other methods in f's class, which stands for the context you must hold in your head to change f safely. A brand-new class has WMC_other = 0, which the max(…, 1) floors to 1, so a new isolated unit still costs something. That floor is what closes the fragmentation loophole. Δlines(f) is the changed-line count in f, and for a new function it is that function's own size. files_changed is the number of distinct production Java files the commit touched, applied as a multiplier, so scattering one rule across many classes is not free. The sum covers functions with no counterpart at the previous commit. For these Δlines(f) is the new function's own line count, and WMC_other is 0 in a brand-new class, floored to 1.

In this harness. Same single pass as the composite.

How to read it. Because of the floor and the spread multiplier, both architectures pay something for additions. So this component does not separate them cleanly, and that is correct. The discrimination is supposed to live in the mutation term.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (composite): purely additive rules only

impact_composite_add  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (composite): purely additive rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that only add a rule and never revise an earlier one. This is the easy case, where an architecture is not asked to change its mind about anything.

How it is calculated
value(c) = impact_composite(c) if c is additive, else blank

Where. A change request is additive when the checkpoint plan declares no mutates list for it. The filtered field is left blank, not zero, on the other checkpoints. Blank matters: a zero would drag the mean of this view toward nothing on every revision checkpoint, which would make the two views incomparable.

In this harness. Derived in the gallery and in analyze from the base field and the checkpoint_type column. It is never written to the per-checkpoint CSV.

How to read it. If a condition only looks good here, it only looks good when nothing has to change.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (composite): rule-revision rules only

impact_composite_mut  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (composite): rule-revision rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that deliberately revise an earlier rule. Fourteen of the sixty change requests are of this kind, and they are placed deliberately rather than at random.

How it is calculated
value(c) = impact_composite(c) if c is mutative, else blank

Where. A change request is mutative when the checkpoint plan declares a mutates list, naming the earlier rules it is allowed to change. Additive checkpoints are blanked, for the same reason as above.

In this harness. Same derivation. Unlike the correctness metrics, the impact score keeps the mutative checkpoints in its base field rather than discounting them, because the context weight makes a mandated revision genuine architectural signal.

How to read it. The interesting case. Revising an existing rule is where an architecture either pays for having isolated the concern or does not. In the control run the concentrated arm paid roughly 36,000 per revision. The distributed arm paid 4,800.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (mutation of existing functions): purely additive rules only

impact_mutation_add  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (mutation of existing functions): purely additive rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that only add a rule and never revise an earlier one. This is the easy case, where an architecture is not asked to change its mind about anything.

How it is calculated
value(c) = impact_mutation(c) if c is additive, else blank

Where. A change request is additive when the checkpoint plan declares no mutates list for it. The filtered field is left blank, not zero, on the other checkpoints. Blank matters: a zero would drag the mean of this view toward nothing on every revision checkpoint, which would make the two views incomparable.

In this harness. Derived in the gallery and in analyze from the base field and the checkpoint_type column. It is never written to the per-checkpoint CSV.

How to read it. If a condition only looks good here, it only looks good when nothing has to change.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (mutation of existing functions): rule-revision rules only

impact_mutation_mut  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (mutation of existing functions): rule-revision rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that deliberately revise an earlier rule. Fourteen of the sixty change requests are of this kind, and they are placed deliberately rather than at random.

How it is calculated
value(c) = impact_mutation(c) if c is mutative, else blank

Where. A change request is mutative when the checkpoint plan declares a mutates list, naming the earlier rules it is allowed to change. Additive checkpoints are blanked, for the same reason as above.

In this harness. Same derivation. Unlike the correctness metrics, the impact score keeps the mutative checkpoints in its base field rather than discounting them, because the context weight makes a mandated revision genuine architectural signal.

How to read it. The interesting case. Revising an existing rule is where an architecture either pays for having isolated the concern or does not. In the control run the concentrated arm paid roughly 36,000 per revision. The distributed arm paid 4,800.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (new-function additions): purely additive rules only

impact_godclass_add  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (new-function additions): purely additive rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that only add a rule and never revise an earlier one. This is the easy case, where an architecture is not asked to change its mind about anything.

How it is calculated
value(c) = impact_godclass(c) if c is additive, else blank

Where. A change request is additive when the checkpoint plan declares no mutates list for it. The filtered field is left blank, not zero, on the other checkpoints. Blank matters: a zero would drag the mean of this view toward nothing on every revision checkpoint, which would make the two views incomparable.

In this harness. Derived in the gallery and in analyze from the base field and the checkpoint_type column. It is never written to the per-checkpoint CSV.

How to read it. If a condition only looks good here, it only looks good when nothing has to change.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (new-function additions): rule-revision rules only

impact_godclass_mut  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (new-function additions): rule-revision rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that deliberately revise an earlier rule. Fourteen of the sixty change requests are of this kind, and they are placed deliberately rather than at random.

How it is calculated
value(c) = impact_godclass(c) if c is mutative, else blank

Where. A change request is mutative when the checkpoint plan declares a mutates list, naming the earlier rules it is allowed to change. Additive checkpoints are blanked, for the same reason as above.

In this harness. Same derivation. Unlike the correctness metrics, the impact score keeps the mutative checkpoints in its base field rather than discounting them, because the context weight makes a mandated revision genuine architectural signal.

How to read it. The interesting case. Revising an existing rule is where an architecture either pays for having isolated the concern or does not. In the control run the concentrated arm paid roughly 36,000 per revision. The distributed arm paid 4,800.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Cohesion: does a class do one thing

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Lack of cohesion, average class (LCOM)

ck_lcom_mean  ·  ↓ lower is better  ·  Chidamber & Kemerer 1994, via the CK tool

Lack of cohesion, average class (LCOM)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Counts the method pairs in a class that share no field, against the pairs that do. A class whose methods all touch the same state is one idea. A class whose methods touch disjoint state is several classes wearing one name. This is the numeric version of what the plain-English cohesion prompt asked for in words.

How it is calculated
LCOM(C) = max(0, |P| − |Q|)
ck_lcom_mean = mean over classes of LCOM(C)

Where. P is the set of method pairs in C whose accessed-field sets are disjoint. Q is the set of pairs that share at least one field. Low is cohesive. The measure is unbounded above and grows roughly with the square of the method count, so a large class is penalised twice: once for incoherence and once for being large.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure. The harness reads CK's per-class LCOM column and takes the mean, and separately the maximum.

How to read it. The natural place to check whether asking the agent for cohesion actually produced it.

Careful. Unbounded and method-count sensitive. Read it with the normalised LCOM* below and with tight class cohesion, which has the opposite sign. Agreement across all three is what makes a cohesion claim safe.

Lack of cohesion, worst class

ck_lcom_max  ·  ↓ lower is better  ·  C&K 1994, via CK

Lack of cohesion, worst class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The least cohesive class in the codebase. Unlike the mean it cannot be diluted by adding cohesive classes, so it answers whether there is a junk-drawer class in here, rather than whether classes are cohesive on average.

How it is calculated
ck_lcom_max = max over classes of LCOM(C)

Where. Same LCOM as above. Because it is unbounded and grows with method count, the worst class is often simply the largest one, so read it beside the class-weight concentration index.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. A step change here means a junk drawer appeared.

Lack of cohesion, normalised (LCOM*)

ck_lcom_star_mean  ·  ↓ lower is better  ·  Henderson-Sellers 1996, via CK

Lack of cohesion, normalised (LCOM*)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

A redesign of LCOM that is bounded roughly to the range 0 to 1 and does not simply grow with the number of methods. Because it is normalised, a difference here is a difference in shape rather than in class size, which is exactly the correction the original LCOM needs.

How it is calculated
LCOM* (C) = ( (1/a) Σj=1..a μ(aj) − m ) / ( 1 − m )

Where. m is the number of methods in C and a the number of fields. μ(aj) is how many of those methods access field j. So the first term is the average number of methods per field. If every method touches every field the numerator is m − m = 0 and the result is 0, meaning perfectly cohesive. If each field is touched by exactly one method the result approaches 1. Low is cohesive. Undefined for a class with fewer than two methods or no fields.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. The version to prefer when comparing arms whose classes differ in size.

Tight class cohesion (TCC)

ck_tcc_mean  ·  ↑ higher is better  ·  Bieman & Kang 1995, via CK

Tight class cohesion (TCC)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The fraction of method pairs in a class that are directly connected through shared field access. It is reported precisely because its direction is inverted relative to LCOM: if a condition improves LCOM and worsens TCC, the improvement is an artefact of one definition rather than a real gain in cohesion.

How it is calculated
TCC(C) = NDC / NP,   NP = m(m − 1) / 2

Where. m is the number of visible methods. NP is therefore every possible pair of them. NDC is the number of pairs that are directly connected, meaning they access at least one instance variable in common. High is cohesive, which is the opposite sign to LCOM. Undefined, and blank, for a class with fewer than two methods.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Agreement with LCOM in the opposite direction is what makes a cohesion claim safe. Disagreement means one definition is doing the work.

Loose class cohesion (LCC)

ck_lcc_mean  ·  ↑ higher is better  ·  Bieman & Kang 1995, via CK

Loose class cohesion (LCC)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same as tight cohesion, but it also counts methods connected indirectly, through a chain of other methods. It is always at least as high as the tight version, and the gap between them is informative on its own.

How it is calculated
LCC(C) = (NDC + NIC) / NP

Where. NIC is the number of pairs connected only indirectly: not sharing a field themselves, but linked through a chain of methods that do. NDC and NP are as in TCC. High is cohesive.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. A large gap between LCC and TCC means the class holds together only through intermediaries, which is weaker cohesion than the LCC figure alone suggests.

Lack of cohesion of the handler class

ck_handler_lcom  ·  ↓ lower is better  ·  C&K 1994, via CK, pinned to the handler class

Lack of cohesion of the handler class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Cohesion of the one class both architectures agree is the entry point. Codebase averages can be moved by adding files. This cannot.

How it is calculated
ck_handler_lcom = LCOM(H)

Where. H is the handler class, matched by name in CK's per-class output. Same LCOM definition as the codebase mean. Low is cohesive.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. The like-for-like cohesion comparison. A controller accumulating rules that touch disjoint state climbs here.

Tight cohesion of the handler class

ck_handler_tcc  ·  ↑ higher is better  ·  Bieman & Kang 1995, via CK, pinned to the handler class

Tight cohesion of the handler class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The inverted-sign cohesion check, on the handler class.

How it is calculated
ck_handler_tcc = TCC(H) = NDC(H) / NP(H)

Where. Same TCC definition. High is cohesive. Undefined, and therefore blank, for a handler class with fewer than two methods, because there are no pairs to connect.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Sparse by construction in whichever arm keeps its handler minimal. The blanks are a result, not missing data: a handler with one method has no cohesion to measure.

Depth of inheritance tree

ck_dit_mean  ·  ↓ lower is better  ·  Chidamber & Kemerer 1994, via CK

Depth of inheritance tree

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How deep the class hierarchy goes, on average. It is reported to show that neither architecture is achieving its structure through inheritance, which matters because none of the other metrics on this page would attribute complexity hidden in a hierarchy correctly.

How it is calculated
DIT(C) = number of edges from C up to the root
ck_dit_mean = mean over classes of DIT(C)

Where. A class extending nothing has DIT = 1 in CK's convention, counting Object as the root. Interfaces implemented do not add depth.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Flat and low in both arms is the expected and desired result. A rise would mean complexity moved into a hierarchy.

The price of distributing

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Call hops from a handling step

indirection_median  ·  ↓ lower is better  ·  harness breadth-first search over the call graph

Call hops from a handling step

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many calls deep you have to go, from a handling step, to reach the code that does the work. Every hop is a file a developer has to open. This is the sceptic's objection made numeric: you did not remove the complexity, you buried it behind indirection.

How it is calculated
depth(k) = length of the shortest call path from any node to k
indirection_median = mediank ∈ reached depth(k)

Where. reached is every method reachable from the handling nodes. The handling nodes themselves have depth 0. Depth is assigned by breadth-first search, so it is the shortest path, not the longest. The harness also records how many methods were reached, as indirection_reached, which is the denominator.

In this harness. placement.indirection_stats, sharing the same call index as the comprehension metrics.

How to read it. Expected to favour the concentrated arm. Everything in one method is zero hops away. This is Brooks' accidental complexity as a number, and it is the honest price of decomposition.

Careful. It understates a pipeline architecture. Depth is measured from the handling nodes, so a pipeline's twenty wired steps are all at depth 0 while a single-handler arm has exactly one depth-0 method. Worse, a pipeline's hops between steps are declared in YAML and dispatched by the container, so they are not Java calls and do not appear here at all. The honest reading of a pipeline arm is this depth plus its step count. Quoting this column alone would let the arm with the most indirection report the least.

Deepest call chain

indirection_max  ·  ↓ lower is better  ·  harness breadth-first search over the call graph

Deepest call chain

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The longest shortest-path from a handling step to any reachable method. It is the worst-case navigation cost for one change.

How it is calculated
indirection_max = maxk ∈ reached depth(k)

Where. Same depth assignment as the median. Because every depth is a shortest path, this is the eccentricity of the reachable set rather than the length of the longest walk.

In this harness. Same single pass.

How to read it. A rising maximum with a flat median means one long tail appeared rather than a general deepening.

Share of reached methods more than two hops away

indirection_deep_share  ·  ↓ lower is better  ·  harness breadth-first search over the call graph

Share of reached methods more than two hops away

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Of everything the request path reaches, what fraction is far enough away that you would not find it by reading the handler. It is a distribution-shape measure rather than an extreme: not how deep the deepest is, but how much of the system lives out in the far field.

How it is calculated
indirection_deep_share = | { k : depth(k) > 2 } | / | reached |

Where. The threshold of 2 is a choice, and it is the only arbitrary constant in this group. It is meant as “further than the handler and the thing it obviously calls”.

In this harness. Same single pass.

How to read it. Expected to be a counter-signal, like the rest of this group.

Coupling between objects, average class

ck_cbo_mean  ·  ↓ lower is better  ·  Chidamber & Kemerer 1994, via CK

Coupling between objects, average class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many other classes a class depends on. Splitting one class into six creates coupling between the six that did not exist before, so this is expected to be a counter-signal.

How it is calculated
CBO(C) = | { D ≠ C : C references D or D references C } |
ck_cbo_mean = mean over classes of CBO(C)

Where. A reference is any use of the other class: a field type, a parameter type, a local variable, a method call, a thrown exception. CK counts distinct classes, not distinct references, so calling one class fifty times still counts 1. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. If a distributed architecture keeps this flat while adding units, that is a real result in its favour, because it is the objection you would most expect to land.

Coupling between objects, worst class

ck_cbo_max  ·  ↓ lower is better  ·  C&K 1994, via CK

Coupling between objects, worst class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The single most entangled class in the codebase.

How it is calculated
ck_cbo_max = max over classes of CBO(C)

Where. Same CBO definition as the mean.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Where the mean is diluted by many small classes, the maximum is not. A rising maximum with a flat mean means one class is becoming the hub.

Fan-out, average class

ck_fanout_mean  ·  ↓ lower is better  ·  CK

Fan-out, average class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many other classes a class calls out to. It is the directional half of coupling: not how entangled a class is, but how much it depends on others.

How it is calculated
fanout(C) = | { D : C references D } |
ck_fanout_mean = mean over classes of fanout(C)

Where. Unlike CBO this counts only outgoing references. CK records the incoming direction separately as fan-in, which the harness also stores.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Rising fan-out with flat fan-in is the orchestrator shape: a class that coordinates rather than one that is depended upon.

Response for a class, average

ck_rfc_mean  ·  ↓ lower is better  ·  Chidamber & Kemerer 1994, via CK

Response for a class, average

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many distinct methods could end up executing in response to one message to this class. That is its own methods plus everything they call. It is roughly the size of the behaviour you have to consider when you call into a class, which makes it both a testability and a comprehension measure.

How it is calculated
RFC(C) = | M(C) ∪ ⋃m ∈ M(C) R(m) |
ck_rfc_mean = mean over classes of RFC(C)

Where. M(C) is the methods declared by C. R(m) is the set of methods invoked by m. CK uses the one-level version, which is the common implementation: it counts methods called directly by the class's own methods and does not recurse.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. It rises with both concentration and indirection, which makes it a useful tiebreaker between those two stories.

Response for a class, worst

ck_rfc_max  ·  ↓ lower is better  ·  C&K 1994, via CK

Response for a class, worst

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The class with the largest response set. The worst thing in the codebase to call into.

How it is calculated
ck_rfc_max = max over classes of RFC(C)

Where. Same one-level RFC definition as the mean.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Read it beside the handler's own RFC below. If the worst class in the codebase is the handler, that is the thesis. If it is something else, say so.

Response set of the handler class

ck_handler_rfc  ·  ↓ lower is better  ·  CK, pinned to the handler class

Response set of the handler class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How much behaviour is reachable from one call into the endpoint's class. It is closely related to the whole-path complexity in the comprehension group, but computed by a different tool, in units of methods rather than branches, and only one level deep.

How it is calculated
ck_handler_rfc = | M(H) ∪ ⋃m ∈ M(H) R(m) |

Where. H is the handler class. One level only, so unlike the comprehension walk this does not follow the call graph transitively.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. A second opinion on the handler's comprehension load, from a tool with a different parser and a different definition.

Coupling of the handler class

ck_handler_cbo  ·  ↓ lower is better  ·  CK, pinned to the handler class

Coupling of the handler class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many other classes the endpoint's own class depends on. A handler that collects dependencies is a handler that is accumulating responsibilities.

How it is calculated
ck_handler_cbo = | { D ≠ H : H references D or D references H } |

Where. Same CBO definition, scoped to the handler class.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. In a pipeline architecture the dependencies move to the wiring, which is exactly the kind of relocation the framework-dispatch caveat is about. A very low figure here for an arm implementing sixty rules is a prompt to check the validity group.

Law of Demeter violations

pmd_demeter_violations  ·  ↓ lower is better  ·  Lieberherr & Holland 1989, via PMD

Law of Demeter violations

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Places where code reaches through one object to get at another, as in a.getB().getC().doThing(). It is the classic symptom of a class knowing too much about its neighbours' internals, and it is another externally defined detector with thresholds nobody here chose.

How it is calculated
pmd_demeter_violations = count of method calls whose receiver is not this, a parameter, a locally created object, or a field of this

Where. The law says a method may only call methods on: itself, its own parameters, objects it created, and its own fields. PMD's rule reports one violation per offending call site, so a single long chain can contribute several.

In this harness. PMD's LawOfDemeter rule, from the same single spawn as the complexity measures.

How to read it. It tends to rise with distribution, so treat it as a counter-signal. It is also a famously noisy rule, so read the trend rather than the absolute count.

Propagation cost

propagation_cost  ·  ↓ lower is better  ·  MacCormack, Rusnak & Baldwin 2006

Propagation cost

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

If you change a random file, what fraction of the codebase could feel it? It is the density of the transitive closure of the file dependency matrix, and it is the one whole-architecture coupling number here with real pedigree in the modularity literature.

How it is calculated
propagation_cost = Σi=1..n | reach(i) | / n2

Where. n is the number of production Java files. reach(i) is the set of other files reachable from file i by following call edges transitively, so a file does not count itself. The numerator is therefore the number of ordered reachable pairs, and dividing by n2 gives the expected fraction of the system a random change can touch.

In this harness. placement.propagation_cost collapses the method call graph to a file graph, dropping self-edges, then runs a depth-first reach from every file.

How to read it. Lower means better modularised. Use the between-arm comparison at the same change request and nothing else.

Careful. Two load-bearing caveats. First, the n2 denominator rewards having more files, so an arm that splits the same code over more files scores lower for free. Second, it inherits a conservative call resolver that drops every edge it cannot prove. MacCormack reports 10 to 60 percent for real systems, so a value an order of magnitude below that means edges are missing rather than that the design is exceptional. Never quote the absolute value.

Files reachable from a typical file

propagation_fanout_median  ·  ↓ lower is better  ·  MacCormack et al. 2006, unnormalised

Files reachable from a typical file

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The raw count behind propagation cost. From one file, how many files can you reach by following dependencies. Because it is not normalised it cannot be improved by adding files, which makes it the honest version.

How it is calculated
propagation_fanout_median = mediani=1..n | reach(i) |

Where. Same file graph and same transitive reach as the propagation cost, with no n2 division. The harness also records the maximum and the file count.

In this harness. Same single pass as the propagation cost.

How to read it. If the normalised and unnormalised numbers tell different stories, the difference is packaging rather than structure. That comparison is the reason both are reported.

Duplication and boilerplate

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Verbosity: duplicated and anti-pattern lines

verbosity  ·  ↓ lower is better  ·  SlopCodeBench (arXiv:2603.24755) Eq. 4

Verbosity: duplicated and anti-pattern lines

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The share of the codebase that is either copy-pasted from elsewhere in the same codebase or matches a known anti-pattern. This is the metric that catches the cheap way to score well on everything else on this page, which is to copy the logic into a new small class instead of factoring it out.

How it is calculated
verbosity = | clone_lines ∪ pattern_lines | / LOC

Where. clone_lines and pattern_lines are sets of (file, line number) pairs, not counts, which is what makes the union meaningful. The union rather than the sum matters: a line that is both duplicated and an anti-pattern is charged once. LOC is the production Java line count. The result can exceed 1 in principle, because the LOC denominator counts function bodies while the line sets are gathered over whole files.

In this harness. metrics.verbosity. Clones come from jscpd, anti-patterns from PMD or ast-grep depending on the run's own config snapshot, so an old run replays with the detector it actually used. If one detector cannot run, the metric is computed from the other and the harness records which halves ran.

How to read it. One of the four conditions produced exactly the failure this metric exists to catch. Read it beside the placement group, not on its own.

Duplicated lines

verbosity_clone_lines  ·  ↓ lower is better  ·  jscpd clone detection

Duplicated lines

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The count of lines that appear as a near-identical block somewhere else in the codebase. It is the copy-paste half of verbosity, reported on its own so the two halves can be told apart.

How it is calculated
verbosity_clone_lines = | clone_lines |

Where. clone_lines is the set of (file, line) pairs jscpd reports as belonging to a duplicated block. Both copies of a clone are counted, because both are lines a maintainer has to keep in step. jscpd's minimum block size is what decides whether a short repeated idiom counts, and it is pinned in tools/package.json so the threshold cannot drift between runs.

In this harness. jscpd at a pinned version, over the arm's source directories. quality_selftest fails closed if the pinned version is not what is installed.

How to read it. A rising line here while the structural metrics improve is the signature of fragmentation masquerading as decomposition.

Anti-pattern lines

verbosity_pattern_lines  ·  ↓ lower is better  ·  PMD ruleset, or the ast-grep rules in astgrep-rules/

Anti-pattern lines

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The count of lines matching structural anti-patterns. The rules are defined over the syntax tree rather than over text, so they survive reformatting and renaming, which makes them harder to game than a text-based lint.

How it is calculated
verbosity_pattern_lines = | pattern_lines |

Where. pattern_lines is the set of (file, line) pairs flagged by the configured detector. With PMD the ruleset is pmd-rules/java-wasteful.xml. With ast-grep it is the rules in astgrep-rules/. Which one applies is read from the run's own config snapshot.

In this harness. From the same single PMD spawn as the complexity measures, filtered to the wasteful ruleset's rule names. The quality gate routes through the same choice, so gate and metric can never disagree about what a smell is.

How to read it. The other half of verbosity. If this moves and the clone count does not, the agent is writing fresh bad code rather than copying old code.

Validity Guards

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Framework-dispatched classes: the call-graph escape counter

container_total  ·  · descriptive  ·  harness; counts advice, aspects, filters, entity listeners and validators

Framework-dispatched classes: the call-graph escape counter

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many classes are invoked by the framework rather than by an ordinary method call. This is not a finding. It is a lie detector for every other metric on this page that follows a call graph.

How it is calculated
container_total = advice + aspect + filter + entity_listener + validator

Where. Each term counts classes carrying the corresponding framework hook: advice is @ControllerAdvice or @RestControllerAdvice; aspect is @Aspect; filter is a servlet Filter or a HandlerInterceptor; entity_listener is @EntityListeners or a @PrePersist callback; validator is a ConstraintValidator. The harness stores each term separately as well as the total.

In this harness. placement.container_dispatch scans the production Java files for those annotations and interfaces at the checkpoint commit.

How to read it. Not a finding. A lie detector. No call-graph walk can see a class the container dispatches, so every comprehension metric on this page is measuring a shrinking fraction of the code wherever this line rises. In one condition, chains ended with eighteen advice classes and a handler containing the stock upstream body. The whole-path complexity read 3 for a codebase implementing all sixty rules. If this line has moved off its baseline for a condition, discount that condition's comprehension numbers rather than the thesis.

Invalid test gates

gate_invalid  ·  ↓ lower is better  ·  harness resilience layer

Invalid test gates

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Checkpoints where the test run itself failed in a way that would otherwise be scored as a mass regression. A crashed Surefire fork is the usual cause. It reports a successful build and zero selected tests, which naively scores as the entire prior suite regressing at once.

How it is calculated
gate_invalid(c) = 1 if build_ok ∧ selected(c) = 0 ∧ neighbours select many, else 0

Where. The signature is the conjunction: the build succeeded, no tests were selected, and the surrounding checkpoints select dozens. A checkpoint flagged this way is excluded from correctness scoring and retried rather than recorded as a failure.

In this harness. Detected in harness/correctness.py, which blanks every correctness field and sets this flag. analyze also repairs old captures by the same signature, which is why any correctness number produced before the fix must be recomputed rather than trusted.

How to read it. Should be flat at zero. It exists because an unflagged crashed fork once read as 143 regressions where the true figure was 0. The biggest risk to a study like this is infrastructure, not statistics.

Tests selected per checkpoint

total_selected  ·  ↑ higher is better  ·  harness test selection

Tests selected per checkpoint

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many black-box acceptance tests ran at this checkpoint. That is this rule's own tests plus every prior rule's. It rises by construction as rules accumulate, and that rise is what makes the later change requests harder than the early ones.

How it is calculated
total_selected(c) = Σj=1..c | tests(j) |

Where. tests(j) is the acceptance tests belonging to change request j, across all four families. Selection is by change request number, so it is deterministic and identical in both arms.

In this harness. The harness selects by checkpoint from acceptance/ and runs them with Surefire after every checkpoint.

How to read it. A smooth rise is correct. A sudden drop to zero at a checkpoint whose neighbours pass dozens is the crashed-fork signature above.

Classes seen by the independent parser

ck_classes  ·  · descriptive  ·  CK tool

Classes seen by the independent parser

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many classes the second parser could read. A parser that fails on a file produces no metrics for it, which makes that file look perfect, so this is the cheapest available check that both tools are seeing the same codebase.

How it is calculated
ck_classes = number of rows in CK's per-class output

Where. CK counts real class declarations, so nested and inner classes appear as their own rows. That is why this figure normally sits above the file count rather than equal to it.

In this harness. Row count of CK's class CSV at the checkpoint commit.

How to read it. Compare the shape of this line against the file count. A divergence means CK started failing on something, and every CK metric on this page is then measuring less than it claims.

Tuesday, 22 September 2026

We asked nicely and it cost 23 percent

A series on how software architecture shapes AI driven code degradation. This is the plain version of a result. The paper has the numbers and the caveats.

Quick recap of the experiment. An AI agent gets sixty change requests, one after another, all landing on the same REST endpoint. Add a validation rule. Change how phone numbers are stored. Add a duplicate check. The full test suite runs after every single one.

We run that on a normal Spring codebase. We run the same sixty changes on the same application built with OfficeFloor. Then we look at the wreckage.

The Spring result has always been the same. The create owner handler grows. Rule after rule lands in the same method. By the end it is the biggest thing in the codebase.

Last time we put our scoring formula straight into the prompt. We told the AI exactly how it was being measured. It optimised the score beautifully. It also scattered the rules into thirty one new classes, wrote more duplicated code, and in two runs hid every rule somewhere our tooling could not see. Worse, it stopped keeping earlier rules working. The whole suite was green after only 44% of changes, down from 79%.

That left one obvious question hanging. Was that because we showed it the metric? Or would any instruction about structure have done the same damage?

So we ran it again. This time we just asked nicely.

The whole intervention is one paragraph

No formula. No mention of any metric. No mention of the experiment. Just the paragraph a senior engineer might add to a ticket.

As you do, keep the code well structured: put each piece of
logic where it belongs, in a small unit with a single
responsibility, and reuse existing code instead of copying
it. Do not let any one class or method grow into a catch-all
that accumulates unrelated logic.

Everything else stayed identical. Same sixty changes. Same tests. Same model. Ten independent runs per codebase. Nothing checked the AI's work and made it try again. The paragraph is the entire change.

It worked

The controller file starts at 203 lines. Here is how many lines got added to it over sixty rules.

Spring controller, lines added over 60 rulesTen runs
Normal prompt604 to 962
Asked nicely36 to 273
Told the formula2 to 44

One paragraph of plain English cut the god method to about a fifth of its size. Ask the codebase which class is heaviest at the end and the answer changes too. Under the normal prompt it is the controller in nine runs out of ten. After the paragraph it is the Owner entity in ten runs out of ten. That class is heavy because it holds getters and setters. It does not hold decisions.

How hard is it to read the create path at the end? We add up the complexity of everything reachable from the endpoint. That number goes from 201 down to 129. Every single run landed below the worst normal run.

So the short version is that asking works. You do not need to show the AI your metric to get most of the benefit of it.

It was not free

This is the part that surprised us, and it is the reason this post has the title it has.

Per run of 60 changes, SpringNormal promptAsked nicely
What the agent cost$78$97
Tool turns it took1,5221,844
Wall clock4.0 hours4.6 hours

The bill went up 23%. Roughly five extra tool calls per change. The OfficeFloor side went up 17%.

Now compare that to the formula version. That one cost nothing extra at all. Same money, same turns.

That contrast is worth sitting with. Optimising arithmetic is cheap. Exercising judgement is not. When you ask an agent to think about where code belongs, it reads more, looks around more and edits more. You pay for all of it.

One thing we checked, because it is the obvious follow up. Within the ten runs, did the ones that spent more end up better structured? No. Spending more did not buy more. The 23% is the price of the instruction. It is not a dial you can turn.

The code still moved rather than shrank

We have a check that ignores all of our metrics. It takes the finished codebase, finds every line the run changed, works out which method that line now sits in, and adds up how complicated all those methods are. It just asks how much logic exists and where it lives.

Spring, after 60 rulesNormal promptAsked nicelyTold the formula
Total logic written282326268
...sitting in brand new files46205218
...sitting in files that already existed23612150
Number of files involved195841
Duplicated lines at the end2,2302,8202,510

Look at the first row. Sixty business rules are sixty business rules. The logic did not get smaller. It got slightly bigger, because small classes need declarations, constructors and call sites.

The next two rows are the same story as last time. The logic moved out of the files that existed and into files the AI created. Three times as many files as the normal run.

Then look at the last row. Duplication went up by about a quarter. Remember the paragraph we wrote. It contains the words "reuse existing code instead of copying it". That is the one explicit request in it. It is the one thing the run did not deliver.

We do not have a proven reason for that. Our best guess is that the other instruction won. Keep units small, and you end up with a lot of small units. Finding the right existing helper among sixty of them is harder than writing a fresh one. If you have ever worked in a codebase with four slightly different StringUtils classes, you have seen a human do the same thing.

The good news, and it is genuinely good

This run created about 41 new classes per run. The formula run created about 31. So more classes. The interesting bit is what kind.

New classes per run, SpringAsked nicelyTold the formula
Proper injected Spring beans167
Static utility holders1621

This matters more than it looks. A static utility class is a class with one static method and nothing else. On our score it is perfect, because there is no surrounding class weight to pay for. In a real Spring codebase it costs you things you run into fast. You cannot inject anything into it. You cannot swap it in a test. Spring cannot wrap it in a transaction or a proxy.

Told the formula, the AI mostly wrote static holders. That is the cheapest possible answer to "make this number small". Asked in English, it wrote real beans about half the time. That is an actual answer to "put this where it belongs".

And nothing vanished this time. Last run, two of the ten Spring runs moved every rule into interceptors that Spring calls for you. Our comprehension metric follows method calls, and nothing calls an interceptor, so those runs simply disappeared from the measurement. That did not happen here. The count of framework invoked classes stayed near the application's baseline of three, with a worst run of eight. Under the formula it averaged thirteen, with a worst run of thirty three. When the numbers got better this time, the code actually got better.

The part that actually matters

Here is the correctness result across all three prompts.

SpringNormalAsked nicelyTold the formula
Implemented the rule it was asked forevery timeevery timeevery time
Whole test suite still green79%62%44%
First rule permanently broken atrule 47rule 30rule 24

The AI always did the job in front of it. All 1200 changes, in this run and in every other one. What it stopped doing was keeping the previous rules working.

And look where the middle column sits. We never mentioned a metric. We asked for tidy code in plain English. Retention still fell hard, and the first permanent breakage still arrived seventeen rules earlier than normal.

That answers the question we ran this for. Roughly half of the damage we blamed on the metric last time was not about the metric at all. It is the cost of restructuring while you are also trying to land a change. Telling the AI the formula then adds a second helping of damage on top, plus the static holders, plus the runs that disappear.

Why does restructuring cost correctness? We do not know for certain yet. The shape of it suggests the agent is doing two jobs in one turn. Land the new rule. Also rearrange the neighbourhood. Both get done. The blast radius of the second job is what breaks rule 19 while you are busy with rule 30.

What to take from this

If you are working with AI tools, here is the practical version.

Asking for structure works. One plain paragraph did most of what our exact scoring formula did. You do not need to invent a metric and feed it to the model. Plain words about single responsibility and putting logic where it belongs are enough to get most of the benefit.

Expect the bill to rise. Structure is work. The agent reads more and edits more. 23% more here. If someone tells you a prompt makes an agent write better code for free, be curious about what they measured.

A quiet metric is not a clean codebase. Our structural numbers all improved. Duplication went up, total code went up, files tripled, and the share of changes leaving the whole suite green fell from 79% to 62%. Only the tests caught that. Keep your acceptance tests, and run all of them, not just the ones for the thing you just changed.

Watch which way the refactoring goes. There is a real difference between "I moved this logic somewhere better" and "I put my new logic where the old code cannot charge me for it". The first one touches existing code. The second one avoids it. In these runs the plain prompt touched more existing code than normal, which is what real refactoring looks like. The formula prompt touched much less, which is what avoidance looks like.

Be suspicious of a pile of static helper classes. They are sometimes right. They are also the cheapest way to make almost any code metric look better, and in Spring they cost you injection, mocking, transactions and proxies.

The instruction has a target. On the Spring codebase, which had a real structural problem, the paragraph helped a lot. On the OfficeFloor codebase, which already spreads rules across wired functions, the same paragraph mostly bought churn. Some of its numbers got slightly worse. A structural instruction is a remedy. It is not a vitamin.

What we are not claiming

This is one paragraph of prompt, one model, one application and one sequence of sixty changes. A different wording would give different results. In particular the duplication finding is a result about our wording, and our wording is the one that asked for reuse.

We also cannot yet tell you how to get the structure without the correctness cost. That is the open problem.

For now the advice is short. Ask for good structure in words. Keep your scoring function to yourself. Budget for both.