Wednesday, 23 September 2026

Where the complexity sits

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Complexity inequality across files (Gini)

ccdist_file_gini  ·  ↓ lower is better  ·  Gini 1912, applied to per-file cyclomatic complexity

Complexity inequality across files (Gini)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Take every file's total complexity and ask how unequally it is distributed. It is the same statistic economists use for income. 0 means every file carries the same complexity. 1 means one file carries it all. This is the headline concentration metric of the whole experiment, because it is the only one of the four that measures shape alone.

How it is calculated
G = Σi=1..n (2i − n − 1) · xi  /  (n · Σx)

Where. The input is a vector of per-unit weights x1 … xn, one per unit, zeros dropped. pi = xi / Σx is unit i's share of the total. For the Gini the weights must be sorted ascending, so i is the rank from lightest to heaviest and the leading factor runs from negative to positive. xi is file i's summed CC over the functions in it. The result is undefined, and blank, for fewer than two files.

In this harness. placement.gini over per-file CC totals at the checkpoint commit. Files with zero complexity are dropped before sorting, because an interface or a constants holder would otherwise read as poverty.

How to read it. This is the one to lead with. Gini is scale-free: it describes the shape of the distribution and is insensitive to how many units it is spread over. An architecture cannot improve its Gini by splitting files into more files of the same shape. So if Gini separates the arms, the distribution really is more unequal.

Complexity concentration across files (HHI)

ccdist_file_hhi  ·  ↓ lower is better  ·  Herfindahl-Hirschman index (Hirschman 1945), over per-file CC shares

Complexity concentration across files (HHI)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The sum of squared shares. It is the standard concentration index from competition regulation, where it decides whether a market is too concentrated. One file holding everything scores 1. A hundred equal files score 0.01. Squaring is what makes it responsive: it is dominated by the largest holder, so it moves sharply when one file runs away.

How it is calculated
HHI = Σi=1..n pi2

Where. The input is a vector of per-unit weights x1 … xn, one per unit, zeros dropped. pi = xi / Σx is unit i's share of the total. The index ranges from 1/n, for perfectly even, up to 1, for everything in one unit. That lower bound is the whole problem with it here, because it depends on n.

In this harness. placement.hhi over per-file CC totals.

How to read it. Sharper than the Gini when one file is running away, because squaring weights the leader.

Careful. Not scale-free. An arm with more units scores lower for free, whatever the shape of its distribution. Always read it beside the Gini, which is scale-free, and beside the unit count. If Gini agrees, the honest claim is that the distribution is more unequal. If only this moves, the honest claim is only that the units are larger.

Share of all complexity in the single heaviest file

ccdist_file_top1  ·  ↓ lower is better  ·  concentration ratio CR1

Share of all complexity in the single heaviest file

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

What fraction of the entire application's complexity lives in its one biggest file. This is the most directly interpretable number on the page. It reads as a sentence: one file in this codebase holds this much of all its decisions.

How it is calculated
CR1 = maxi xi / Σx

Where. The input is a vector of per-unit weights x1 … xn, one per unit, zeros dropped. pi = xi / Σx is unit i's share of the total. No sorting or squaring. It is simply the largest single share.

In this harness. placement.top_share with k=1, over per-file CC totals.

How to read it. If this climbs toward a third as rules accumulate, the god-file story is not a metaphor. It is also the easiest figure to quote to someone who does not want a statistics lesson.

Share of all complexity in the heaviest five files

ccdist_file_top5  ·  ↓ lower is better  ·  concentration ratio CR5

Share of all complexity in the heaviest five files

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same question widened to five files, so a codebase cannot look healthy simply by splitting its god file in two.

How it is calculated
CR5 = Σi=1..5 x(i) / Σx

Where. The input is a vector of per-unit weights x1 … xn, one per unit, zeros dropped. pi = xi / Σx is unit i's share of the total. x(i) denotes the weights sorted descending, so the sum is over the five heaviest. With fewer than five non-zero files the value is 1 by construction.

In this harness. placement.top_share with k=5.

How to read it. Resistant to the cosmetic split. If the top-1 share falls but this does not, the complexity was moved next door rather than distributed.

Complexity spread across files (normalised entropy)

ccdist_file_hnorm  ·  ↑ higher is better  ·  Shannon entropy, normalised by log n

Complexity spread across files (normalised entropy)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How evenly complexity is spread, on a scale of 0 to 1 where 1 is perfectly even. It is the information-theoretic mirror of the concentration indices, and it reads as the number of bits you would need to say which file a randomly chosen unit of complexity came from.

How it is calculated
H = −Σi pi ln pi
Hnorm = H / ln n

Where. The input is a vector of per-unit weights x1 … xn, one per unit, zeros dropped. pi = xi / Σx is unit i's share of the total. n is the number of non-zero units. Dividing by ln n is the maximum entropy that many units can have, which is what puts the result on a 0 to 1 scale and removes most of the unit count advantage that HHI gives away. The base of the logarithm cancels in the ratio, so natural log here and log base 2 elsewhere give the same normalised number.

In this harness. placement.norm_entropy over per-file CC totals. Undefined, and blank, for fewer than two files.

How to read it. Note the direction flips. Higher means more distributed. Because it is normalised it makes a useful referee between the Gini and the HHI when those two disagree.

Complexity inequality across functions (Gini)

ccdist_fn_gini  ·  ↓ lower is better  ·  Gini 1912, applied to per-function CC

Complexity inequality across functions (Gini)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same inequality question one level down, across individual functions rather than files. Files can be split for cosmetic reasons. A function is the unit a developer actually reads at one sitting, so this is the god-method question rather than the god-file one.

How it is calculated
G = Σi=1..n (2i − n − 1) · xi  /  (n · Σx)

Where. Same formula as the file Gini, with xi now the cyclomatic complexity of one function, sorted ascending. n is the number of functions with non-zero complexity.

In this harness. placement.gini over the per-function CC column.

How to read it. This asks whether one method is becoming the place decisions go. It is harder to game than the file version, because you cannot split a method without changing the code.

Complexity concentration across functions (HHI)

ccdist_fn_hhi  ·  ↓ lower is better  ·  Herfindahl-Hirschman index over per-function CC shares

Complexity concentration across functions (HHI)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Concentration of decisions into single methods. The god-method signal, as opposed to the god-class signal.

How it is calculated
HHI = Σi=1..n pi2

Where. Same formula as the file HHI, with pi now one function's share of total cyclomatic complexity.

In this harness. placement.hhi over the per-function CC column.

How to read it. Squaring means one very heavy method dominates the figure, which is the behaviour you want from a god-method detector.

Careful. Not scale-free. An arm with more units scores lower for free, whatever the shape of its distribution. Always read it beside the Gini, which is scale-free, and beside the unit count. If Gini agrees, the honest claim is that the distribution is more unequal. If only this moves, the honest claim is only that the units are larger.

Complexity concentration across packages (HHI)

ccdist_pkg_hhi  ·  ↓ lower is better  ·  Herfindahl-Hirschman index over per-package CC shares

Complexity concentration across packages (HHI)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Concentration at the coarsest level, across whole packages. It survives any amount of file-level reshuffling inside a package.

How it is calculated
HHI = Σi=1..n pi2

Where. pi is package i's share of total cyclomatic complexity, where a package is the directory path below the source root. n is the package count, which is small, so this index sits much higher than the file or function versions and must not be compared against them numerically.

In this harness. placement.hhi over CC grouped by placement._package_of.

How to read it. If a condition improves the file numbers but not this one, the rules were split into more files inside the same package. That is a smaller change than it looks.

Class-weight concentration (HHI over WMC)

wmcdist_class_hhi  ·  ↓ lower is better  ·  Herfindahl-Hirschman index over per-class WMC, via CK

Class-weight concentration (HHI over WMC)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The concentration question over classes, weighted by how much complexity each class's methods carry. It is the class-level view of the god-object question, computed by the independent parser.

How it is calculated
HHI = Σi pi2,   pi = WMC(Ci) / Σj WMC(Cj)

Where. WMC(C) is the sum of cyclomatic complexity over the methods of class C, as CK reports it. Unlike the file-based indices this one sees nested and inner classes as their own units, because CK resolves real class declarations rather than assuming one class per file.

In this harness. placement.hhi over CK's per-class WMC column.

How to read it. A useful cross-check on the file-level indices, since it uses a different notion of a unit and a different parser.

Careful. Not scale-free. An arm with more units scores lower for free, whatever the shape of its distribution. Always read it beside the Gini, which is scale-free, and beside the unit count. If Gini agrees, the honest claim is that the distribution is more unequal. If only this moves, the honest claim is only that the units are larger.

Cognitive-complexity concentration across functions

cogdist_fn_hhi  ·  ↓ lower is better  ·  HHI over per-function cognitive complexity (PMD)

Cognitive-complexity concentration across functions

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Concentration measured in units of human difficulty rather than branch count. It asks whether the nesting is pooling in one place, which is a different question from whether the branches are.

How it is calculated
HHI = Σi pi2,   pi = cognitive(fi) / Σj cognitive(fj)

Where. cognitive(f) is the Campbell measure defined in the amount group: one point per flow-breaking structure plus one per level of nesting it sits inside.

In this harness. placement.hhi over PMD's per-method cognitive complexity.

How to read it. If concentration shows up in cyclomatic terms but not cognitive terms, the concentrated method is long but flat. If it shows up in both, it is long and deeply nested. That is the genuinely hard kind.

Careful. Not scale-free. An arm with more units scores lower for free, whatever the shape of its distribution. Always read it beside the Gini, which is scale-free, and beside the unit count. If Gini agrees, the honest claim is that the distribution is more unequal. If only this moves, the honest claim is only that the units are larger.

Halstead-volume concentration across files

voldist_file_hhi  ·  ↓ lower is better  ·  HHI over per-file Halstead volume

Halstead-volume concentration across files

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Concentration measured by vocabulary rather than by control flow. It is the control-flow-free confirmation that the concentration result is not an artefact of how branches are counted.

How it is calculated
HHI = Σi pi2,   pi = V(filei) / Σj V(filej)

Where. V(file) is that file's Halstead volume, as defined in the amount group. Because literals are collapsed to one placeholder operand, this cannot be moved by editing message text.

In this harness. placement.hhi over the harness's per-file Halstead volumes.

How to read it. A branch-free second opinion on the file-level concentration result.

Careful. Not scale-free. An arm with more units scores lower for free, whatever the shape of its distribution. Always read it beside the Gini, which is scale-free, and beside the unit count. If Gini agrees, the honest claim is that the distribution is more unequal. If only this moves, the honest claim is only that the units are larger.

Heaviest class in the codebase (WMC)

wmc_max  ·  ↓ lower is better  ·  Chidamber & Kemerer 1994, via lizard

Heaviest class in the codebase (WMC)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Weighted methods per class is the sum of the cyclomatic complexities of every method in a class. The heaviest such class is the god-class candidate. It catches what a per-method erosion threshold cannot: a controller that stays tidy method by method, while accumulating twenty rules' worth of methods, becomes a god class without any single method ever crossing CC 10.

How it is calculated
WMC(C) = Σm ∈ methods(C) CC(m)
wmc_max = maxC ∈ touched WMC(C)

Where. C ranges over the touched subsystem, which is the production Java files changed since the baseline commit. That scoping is deliberate: it tracks what the agent is building rather than what shipped with the framework. On the lizard side a class is identified with a file, on the assumption of one top-level class per Java file.

In this harness. metrics.wmc_stats groups lizard's functions by file, sums CC per group and returns the worst, along with its name, method count and line count.

How to read it. Rising means one class is accumulating everything.

Careful. Role-blind, and that matters here. The two architectures answer this question with different kinds of class. One arm's heaviest is usually the controller, around 32 methods averaging CC above 3. The other's is often the Owner entity, around 60 accessors at CC 1. Those are not the same problem. Use the handler-scoped version below for any claim about the architectures.

Weight of the handler class (like-for-like)

wmc_handler  ·  ↓ lower is better  ·  C&K WMC, pinned to the class the endpoint routes through

Weight of the handler class (like-for-like)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same god-class number, but measured on the class that actually handles the create-owner request, in both architectures. Pinning it to a role rather than to whichever class happens to be heaviest is what makes it a fair comparison.

How it is calculated
wmc_handler = Σm ∈ methods(H) CC(m)

Where. H is the handler class, which is the file containing the function the create endpoint routes through. The harness finds it per arm from the configured entry_handler pattern, so it resolves to the Spring controller in one arm and to the wired create function's class in the other. The field is blank until that class exists.

In this harness. Same scoping as handler erosion, in metrics.entry_handler_stats and its WMC companion.

How to read it. A rising line means the front door is turning into a god class. This is the like-for-like god-class comparison, same role, both arms.

Careful. Not prompt-robust. This can be driven to zero by telling the agent the formula. The logic simply moves into new files while the total complexity is unchanged. It stays valid for the ungated control. Never publish it as the headline for an intervention arm without the whole-path number from the comprehension group beside it.

Weight of the handler class (independent parser)

ck_handler_wmc  ·  ↓ lower is better  ·  CK tool (Aniche 2015), Eclipse JDT parser

Weight of the handler class (independent parser)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The handler-class weight again, from the second tool. A cross-check that neither parser is silently failing on the handler file, which is the one file on which the whole concentration argument rests.

How it is calculated
WMC(H) = Σm ∈ methods(H) CC(m), as CK reports it

Where. H is matched by class name rather than by file path here, because CK reports per class. Differences from the lizard figure are usually nested classes, which CK attributes separately.

In this harness. Read from CK's per-class CSV, filtered to the handler class name.

How to read it. Agreement with the lizard figure is the boring result you want.

Complexity of the front-door method

entry_cc  ·  ↓ lower is better  ·  McCabe 1976, scoped to the entry function

Complexity of the front-door method

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The cyclomatic complexity of the single method the HTTP request lands in. That is the method a developer opens first when asked to change this endpoint, so it is the purest available measure of whether the front door bloats.

How it is calculated
entry_cc = CC(e)

Where. e is the one function the create endpoint routes through, which is addOwner in the Spring arm and the wired create function's service method in the OfficeFloor arm. CC(f) is the cyclomatic complexity of function f, as reported by lizard: one plus the number of decision points, counting if, for, while, case, catch, the ternary operator, and each && or ||.

In this harness. metrics.entry_handler_stats matches the configured entry pattern against lizard's function list and reports that function's CC, line count and name.

How to read it. The purest front-door measure, and the easiest to explain.

Careful. Entry-scoped, so it understates a pipeline architecture by construction. An architecture whose entry node just dispatches will read about 1 while its worst downstream step reads about 13. Never publish this without the whole-path complexity beside it. A reader who opens the wiring file will make that objection for you.

Worst single function in the touched subsystem

hotspot_cc  ·  ↓ lower is better  ·  lizard

Worst single function in the touched subsystem

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The highest cyclomatic complexity of any one function among the files this run has changed. Scoping it to changed files means it tracks what the agent is actually building, rather than whatever the framework shipped with.

How it is calculated
hotspot_cc = maxf ∈ touched CC(f)

Where. touched is the set of production Java functions in files changed since the baseline commit. The harness also records which function it was, as hotspot_fn, so the figure can be traced to a name.

In this harness. metrics.hotspot_stats over the dynamically scoped subsystem.

How to read it. A climbing line means the worst thing the agent has written keeps getting worse. Because it is a maximum it is noisy, so read the trend.

Structural erosion, whole application

erosion  ·  ↓ lower is better  ·  SlopCodeBench (arXiv:2603.24755) Eq. 3

Structural erosion, whole application

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

What fraction of the codebase's complexity sits inside functions that are over the complexity threshold. Each function is given a mass that combines its branching with its size, and erosion is the share of that mass held by functions judged too complex. It is the benchmark's own headline measure, which is why it is reported here at all.

How it is calculated
mass(f) = CC(f) · √max(SLOC(f), 1)
erosion = Σf : CC(f) > 10 mass(f)  /  Σf ∈ F mass(f)

Where. CC(f) is the cyclomatic complexity of function f, as reported by lizard: one plus the number of decision points, counting if, for, while, case, catch, the ternary operator, and each && or ||. SLOC(f) is the function's non-comment line count. The square root is the benchmark's choice: it means size matters but with diminishing returns, so a 400-line function is not simply scored as ten times a 40-line one. The threshold of CC > 10 is McCabe's own recommended limit, not a value chosen here. The population is production Java only, meaning the files matched by the arm's source_globs at that checkpoint's commit. Tests are excluded. OfficeFloor's YAML wiring is counted in a separate pool and is never folded into a Java denominator, because mixing them would let a wiring-based architecture dilute any per-line metric.

In this harness. metrics.erosion_detail. The harness records the numerator and denominator separately, as erosion_high_mass and erosion_total_mass, so the ratio is reproducible from the CSV without rerunning anything.

How to read it. Kept for comparability with the published benchmark.

Careful. Demoted from decisive in this experiment. It is location-blind. It cannot tell a CC-19 god method from a CC-19 isolated single-responsibility algorithm. Here it is also dominated by architecture-neutral leaf algorithms (soundex, phone normalisation, deduplication) that both arms have to implement. Its arm ordering has come out backwards. Read it as a leaf-algorithm measure, not as a test of the thesis.

Structural erosion, changed files only

erosion_scoped  ·  ↓ lower is better  ·  SlopCodeBench Eq. 3, scoped to the touched subsystem

Structural erosion, changed files only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same erosion ratio, restricted to files this run has actually modified, so untouched framework code cannot dilute it.

How it is calculated
erosion_scoped = Σf ∈ touched, CC(f) > 10 mass(f)  /  Σf ∈ touched mass(f)

Where. Same mass and same threshold as the whole-app version. touched is the production Java functions in files changed since the baseline commit, which grows as the run proceeds. The harness records the subsystem's function count as subsystem_nfns so the denominator's size is visible.

In this harness. metrics.erosion_detail over the dynamically scoped subsystem.

How to read it. A tighter version of the whole-app number.

Careful. It inherits the same location-blindness. A big isolated algorithm in a changed file scores exactly the same as a big god method there.

Structural erosion of the handler class

erosion_handler  ·  ↓ lower is better  ·  SlopCodeBench Eq. 3, scoped to the entry handler's own class

Structural erosion of the handler class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Erosion measured only inside the class that handles the request. Scoping it this way excludes the leaf algorithms that swamp the two wider versions. What is left is the clean concentration signal: does the endpoint's own handler surface rot as rules accumulate?

How it is calculated
erosion_handler = Σf ∈ H, CC(f) > 10 mass(f)  /  Σf ∈ H mass(f)

Where. H is the functions of the handler class only, found by the same entry pattern as the handler WMC. The harness records erosion_handler_class and erosion_handler_nfns so you can confirm which class was measured and how many functions were in the denominator.

In this harness. metrics.handler_scoped_erosion.

How to read it. This is the erosion number that tests the thesis.

Careful. Like the handler WMC, this can be driven to exactly zero by an intervention that relocates logic without removing it. It stays valid for the ungated control. For an intervention arm, read the comprehension group instead.

God classes detected

pmd_god_classes  ·  ↓ lower is better  ·  Lanza & Marinescu 2006 thresholds, via PMD's GodClass rule

God classes detected

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many classes trip a published god-class detector. The value of this metric is that nobody in this experiment chose the thresholds. It is an externally defined, binary verdict, which makes it immune to the accusation that the scoring was tuned to the result.

How it is calculated
GodClass(C) ⇔ WMC(C) ≥ 47  ∧  ATFD(C) > 5  ∧  TCC(C) < 1/3
pmd_god_classes = | { C : GodClass(C) } |

Where. All three conditions must hold. WMC is weighted methods per class, as above. ATFD is Access To Foreign Data: the number of distinct fields of other classes this one reaches into, which is what distinguishes a class doing too much from a class that is merely large. TCC is tight class cohesion, defined in the cohesion group, so the third condition says the class is also internally incoherent. The constants 47, 5 and 1/3 are Lanza and Marinescu's.

In this harness. PMD's GodClass rule, from the same single PMD spawn as the complexity measures. The harness counts distinct violating classes.

How to read it. A step change in this line is a strong, externally validated statement that the codebase acquired a god class at that change request.

Is the handler class itself a god class

pmd_handler_is_god_class  ·  ↓ lower is better  ·  PMD GodClass rule, evaluated on the handler class

Is the handler class itself a god class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Whether the class handling the endpoint trips the published detector. It is binary, externally defined, and about the specific class both architectures agree is the entry point, which makes it the single most quotable structural result on this page.

How it is calculated
value(c) = 1 if GodClass(H) at change request c, else 0
plotted value = (1/K) Σk=1..K valuek(c)

Where. H is the handler class. GodClass is the three-condition test above. K is 10, the number of runs, so the plotted line is the fraction of runs in which the front door has become a god class by that change request.

In this harness. The harness passes the handler class name to placement.pmd_metrics_from_report and records whether PMD flagged it.

How to read it. Read the line as: in what fraction of runs has the front door become a god class by change request N.

Data classes detected

pmd_data_classes  ·  · descriptive  ·  PMD DataClass rule

Data classes detected

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Classes that are mostly fields and accessors with little behaviour. This is context rather than a finding. It is how you tell whether a heaviest-class result is about a controller full of decisions or an entity full of getters.

How it is calculated
pmd_data_classes = | { C : DataClass(C) } |

Where. PMD's DataClass rule fires on a class with a high proportion of public accessors, few real methods and low complexity per method. It is the complement of the god class: too little behaviour rather than too much.

In this harness. Same single PMD spawn, counting distinct violating classes.

How to read it. Use it to interpret the role-blind heaviest-class metric. A rising data class count in one arm usually means its entity is growing accessors.

What you must understand to change one rule

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Complexity reachable from a typical handling step

node_cc_median  ·  ↓ lower is better  ·  harness call-graph walk from the declared wiring nodes

Complexity reachable from a typical handling step

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Pick one step in the request's handling. Follow every method it calls, transitively, and add up the complexity. That total is what a developer must understand to change that one step. This metric is the median of that total across all the steps. It is the closest thing on this page to the real question, which is what it costs to change one rule.

How it is calculated
node_cc_median = medianr ∈ nodes cc(closure(r))

Where. Nodes are the steps the request passes through. They come from the arm's own declared wiring, which is the YAML file named by node_roots.wiring_file. An arm that declares no wiring contributes exactly one node, its entry_handler. closure(r) is the set of methods reachable from node r by following Java call edges transitively, computed by breadth-first search over a call graph the harness builds from lizard's function list plus name resolution. cc(S) = Σk ∈ S CC(k) for a set of methods S.

In this harness. metrics.node_closure_stats. The call index is built once per checkpoint and shared with the indirection and propagation metrics, because resolving the graph is the expensive part.

How to read it. This is the concentration statistic to lead with. It is relocation-proof: work pushed into a helper still lands in that helper's caller's closure, so moving code downstream does not improve it. It is the metric that survived the condition where the agent was told the scoring formula.

Careful. Code the framework dispatches is not code the handler calls. A rule moved into a @RestControllerAdvice, an @Aspect, a servlet Filter, a @PrePersist entity listener or a ConstraintValidator leaves this walk entirely, because the container invokes it and there is no Java call edge to follow. Check the framework-dispatch counter in the validity group before trusting this for a given condition.

Complexity reachable from the worst handling step

node_cc_max  ·  ↓ lower is better  ·  harness call-graph walk

Complexity reachable from the worst handling step

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same closure complexity, for whichever step is worst. The median says what a typical change costs. This says what the worst change costs, which is often what a team actually remembers.

How it is calculated
node_cc_max = maxr ∈ nodes cc(closure(r))

Where. Nodes are the steps the request passes through. They come from the arm's own declared wiring, which is the YAML file named by node_roots.wiring_file. An arm that declares no wiring contributes exactly one node, its entry_handler. closure(r) is the set of methods reachable from node r by following Java call edges transitively, computed by breadth-first search over a call graph the harness builds from lizard's function list plus name resolution. cc(S) = Σk ∈ S CC(k) for a set of methods S. The harness also records the mean and the 90th percentile of the same distribution, as node_cc_mean and node_cc_p90.

In this harness. Same single pass as the median.

How to read it. A distributed architecture is allowed a high median and a low maximum, because it has many small steps. It is in trouble if its maximum approaches the concentrated arm's, because that means one of its steps has become the god method it was supposed to avoid.

Careful. Code the framework dispatches is not code the handler calls. A rule moved into a @RestControllerAdvice, an @Aspect, a servlet Filter, a @PrePersist entity listener or a ConstraintValidator leaves this walk entirely, because the container invokes it and there is no Java call edge to follow. Check the framework-dispatch counter in the validity group before trusting this for a given condition.

Complexity of the entire handling path

node_path_cc  ·  ↓ lower is better  ·  harness call-graph walk, unioned over all nodes

Complexity of the entire handling path

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Everything reachable from any step of the request path, counted once. This is the honest total: what the whole feature costs to understand. Taking the union rather than the sum matters, because a shared helper reachable from six steps is one thing to learn, not six.

How it is calculated
node_path_cc = cc( ⋃r ∈ nodes closure(r) )

Where. Nodes are the steps the request passes through. They come from the arm's own declared wiring, which is the YAML file named by node_roots.wiring_file. An arm that declares no wiring contributes exactly one node, its entry_handler. closure(r) is the set of methods reachable from node r by following Java call edges transitively, computed by breadth-first search over a call graph the harness builds from lizard's function list plus name resolution. cc(S) = Σk ∈ S CC(k) for a set of methods S. The union is over method identities, so a helper reached from several nodes is counted exactly once. The harness also records the size of that union as node_path_methods.

In this harness. Same single pass. The set union is taken before summing complexity, not after.

How to read it. This is the answer to “you just moved it downstream”. It must be published beside the front-door and handler-class metrics. It is also where the two architectures have historically come out closest. In one control run it read 229 against 202. That is the same total work, arranged differently. It is Tesler's conservation showing up in the place where it is hardest to argue with.

Careful. Code the framework dispatches is not code the handler calls. A rule moved into a @RestControllerAdvice, an @Aspect, a servlet Filter, a @PrePersist entity listener or a ConstraintValidator leaves this walk entirely, because the container invokes it and there is no Java call edge to follow. Check the framework-dispatch counter in the validity group before trusting this for a given condition. A chain whose rules moved into advice classes can report a path complexity in single digits while implementing all sixty rules. A number that low is a detector of the escape, not a result.

Number of handling steps

node_count  ·  · descriptive  ·  harness, from the architecture's declared wiring

Number of handling steps

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many distinct steps the request passes through. This is the mechanism, not a finding. It is also the correct denominator when reading the median closure complexity, because a rising median across a rising node count is a different story from a rising median at a fixed one.

How it is calculated
node_count = | nodes |

Where. Nodes declared in the arm's wiring file, or 1 for an architecture that declares none. The asymmetry is real and is the point: one arm has a single node by design.

In this harness. metrics._node_roots parses the wiring file and resolves each declared step to a method in the call index. A step it cannot resolve is dropped, so this is a floor rather than a declaration count.

How to read it. Read it beside the median. It is also the honest companion to the indirection metrics, whose depth figure understates a pipeline exactly because all its steps sit at depth zero.

How much of the path belongs to exactly one step

node_exclusive_share  ·  ↑ higher is better  ·  harness call-graph walk

How much of the path belongs to exactly one step

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Of all the complexity reachable from the handling path, what fraction is reachable from only one step. High means each step owns its own logic. Low means the steps are thin wrappers over a shared blob. This is the metric that would expose a fake decomposition, because twenty wired steps that all call the same helper would show a high step count and a low exclusive share.

How it is calculated
reach(k) = | { r : k ∈ closure(r) } |
node_exclusive_share = Σr cc({ k ∈ closure(r) : reach(k) = 1 })  /  Σr cc(closure(r))

Where. Nodes are the steps the request passes through. They come from the arm's own declared wiring, which is the YAML file named by node_roots.wiring_file. An arm that declares no wiring contributes exactly one node, its entry_handler. closure(r) is the set of methods reachable from node r by following Java call edges transitively, computed by breadth-first search over a call graph the harness builds from lizard's function list plus name resolution. cc(S) = Σk ∈ S CC(k) for a set of methods S. reach(k) counts how many nodes can reach method k. Note that the denominator is the sum over nodes, not the union, so a method shared by six nodes is counted six times below the line and zero times above it. That is what makes sharing expensive in this ratio.

In this harness. Same single pass. Blank for an arm with one node, where it would be trivially 1.0. That blank is deliberate: in an arm-versus-arm table a trivial 1.0 would read as perfect cohesion when it actually means there are no separable rules to share between.

How to read it. This is the cohesion test for a pipeline architecture. It is the number to ask for when someone claims a decomposition is only cosmetic.

Methods reachable from a typical handling step

node_methods_median  ·  ↓ lower is better  ·  harness call-graph walk

Methods reachable from a typical handling step

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same closure, counted in methods rather than in complexity. It is how many distinct methods you would have to read. Because it weights every method equally it is immune to any argument about how complexity should be scored.

How it is calculated
node_methods_median = medianr ∈ nodes | closure(r) |

Where. Nodes are the steps the request passes through. They come from the arm's own declared wiring, which is the YAML file named by node_roots.wiring_file. An arm that declares no wiring contributes exactly one node, its entry_handler. closure(r) is the set of methods reachable from node r by following Java call edges transitively, computed by breadth-first search over a call graph the harness builds from lizard's function list plus name resolution. cc(S) = Σk ∈ S CC(k) for a set of methods S.

In this harness. Same single pass as the complexity median.

How to read it. A count-based confirmation of the closure result. If it agrees with the complexity median, the finding does not depend on McCabe weights.

Careful. Code the framework dispatches is not code the handler calls. A rule moved into a @RestControllerAdvice, an @Aspect, a servlet Filter, a @PrePersist entity listener or a ConstraintValidator leaves this walk entirely, because the container invokes it and there is no Java call edge to follow. Check the framework-dispatch counter in the validity group before trusting this for a given condition.

How much existing code each rule disturbs

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Pre-existing functions modified per rule

existing_fns_modified  ·  ↓ lower is better  ·  harness diff analysis

Pre-existing functions modified per rule

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many functions that already existed, and already worked, had to be opened to land this one change request. This is the blast radius proper. It is the measure that most directly explains the regression results, because you cannot break what you did not edit.

How it is calculated
existing_fns_modified(c) = | { f : file(f) existed at c−1 ∧ lines(f) ∩ changed(c) ≠ ∅ } |

Where. changed(c) is the set of line ranges the checkpoint's diff touched on the new side. lines(f) is function f's line span at the new commit. A function counts if the diff overlapped it at all, including by a single inserted line. Functions in files created by this checkpoint are excluded, because they did not previously exist.

In this harness. metrics.blast_radius_detail takes git diff --name-status -M between the previous and current checkpoint commits, maps changed hunks to functions at the new commit, and counts those in pre-existing files.

How to read it. Zero means the rule was purely additive and nothing that already worked was put at risk. A rising line means the opposite.

New production files created per rule

files_created  ·  · descriptive  ·  harness diff analysis

New production files created per rule

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many new production Java files the change request produced. It has no good direction. It is simultaneously the mechanism by which blast radius stays near zero and the mechanism by which a codebase fragments.

How it is calculated
files_created(c) = | { p : status(p) = A ∧ p is production Java } |

Where. status(p) = A is git's added status for path p in the checkpoint diff, with rename detection on, so a renamed file is not miscounted as a creation. Test files and YAML are excluded.

In this harness. metrics.blast_radius_detail, from the same --name-status -M parse.

How to read it. Read it with the duplication group. New files that copy each other are not a win, and one condition in this study produced exactly that.

Files touched per rule

files_touched  ·  ↓ lower is better  ·  git diff shortstat

Files touched per rule

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many files the change request touched in total, new and existing. The raw spread of one change. High values mean a rule could not be expressed in one place.

How it is calculated
files_touched(c) = | { p : p appears in diff(c−1, c) } |

Where. Every path in the checkpoint diff, before the production-Java filter, so it includes configuration and resources. The harness records the added and removed line counts alongside, as diff_added and diff_removed.

In this harness. git diff --shortstat between consecutive checkpoint commits.

How to read it. The coarsest blast measure, and the one that needs no parser at all, which makes it a useful sanity check on the parsed ones.

Packages touched per rule

packages_touched  ·  ↓ lower is better  ·  harness diff analysis

Packages touched per rule

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many distinct Java packages one change request reached into. It ignores cosmetic file splits inside a package, so it is a coarser and more meaningful spread measure than the file count. A rule that touches four packages is a rule that did not have a home.

How it is calculated
packages_touched(c) = | { package(p) : p ∈ diff(c), p is production Java } |

Where. package(p) is the directory path below the source root, which is the package a conventionally laid out Java file declares.

In this harness. Derived from the production-Java paths in the checkpoint diff.

How to read it. Low and flat is the signature of a rule that had an obvious place to go.

Temporal coupling: how much of this edit was someone else's rule

reedit_rate  ·  ↓ lower is better  ·  harness line-authorship analysis (git blame)

Temporal coupling: how much of this edit was someone else's rule

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Of the lines inside the functions this change request edited, what share was written by earlier change requests. It is the clearest operational statement of the phrase “the rules are tangled”. A high rate means implementing rule 47 required reading and rewriting the code for rules 12 and 30.

How it is calculated
reedit_rate(c) = prior_lines / body_lines
body_lines = Σf ∈ edited(c) | lines(f) |

Where. edited(c) is the pre-existing functions this checkpoint touched. body_lines counts whole function bodies, not just the changed lines. prior_lines is how many of those lines git blame attributes to a commit that is neither this checkpoint nor an ancestor of the baseline, which is exactly the set of earlier change requests. Blank when the checkpoint edited no existing function body.

In this harness. metrics.reedit_stats blames each edited function's body at the current commit and bins every line into one of three eras: the original application, an earlier checkpoint, or this checkpoint. Counting whole bodies is deliberate, because it catches a one-line insertion into a large shared method, which blaming only the diff lines would miss.

How to read it. This is both the comprehension cost and the mechanism for unintended regressions, in one number.

Careful. Blank on any checkpoint that only added new units, which is why the line is sparser in the distributed arm. A blank is a result, not missing data: it means nothing old was reopened.

Change spread for this rule (entropy)

change_entropy_norm  ·  · descriptive  ·  Hassan 2009 change entropy, normalised

Change spread for this rule (entropy)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How evenly this one change request's diff spread across files, on a scale of 0 to 1. Hassan's original finding was that scattered changes predict faults better than the volume of change does, which is why the measure exists at all. It is the cleanest placement measure in the suite, because it is pure git with no parse and nothing bespoke.

How it is calculated
H = −Σi pi log2 pi
Hnorm = H / log2 n

Where. pi is file i's share of the lines this checkpoint changed, taken from git diff --numstat. n is the number of files with at least one changed line. With fewer than two such files the value is 0 by definition, because a change confined to one file has no spread.

In this harness. placement.change_entropy with the previous checkpoint as the left-hand side, over the src/main pathspec.

How to read it. Read this one carefully. The direction depends on your theory. High entropy means the change was scattered, which Hassan associates with faults. But a distributed architecture scatters by design, into files that did not previously exist. Use the cumulative version below for the architectural claim. Use this one for the per-change fault risk.

Share of this rule's changed lines in one file

change_top1  ·  · descriptive  ·  concentration ratio CR1 over the checkpoint diff

Share of this rule's changed lines in one file

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Of the lines this change request touched, what fraction landed in a single file.

How it is calculated
CR1 = maxi linesi / Σj linesj

Where. linesi is the changed-line count in file i, from git diff --numstat. Added and removed lines are both counted.

In this harness. Same numstat parse as the change entropy.

How to read it. Near 1.0 means the whole rule went into one file. For a rule landing in a new file that is ideal. For a rule landing in the same file as the last forty rules it is the god-file mechanism in action. The cumulative metrics below are what distinguish the two cases.

Cumulative change spread (entropy)

cum_change_entropy_norm  ·  ↑ higher is better  ·  Hassan 2009 change entropy, cumulative from the baseline

Cumulative change spread (entropy)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same spread question asked over all the change so far, rather than just this rule. This is the version that answers the architectural question, because it cannot be satisfied by a series of individually tidy diffs that all land in the same place.

How it is calculated
Hnorm = H / log2 n, computed over diff(base, c)

Where. Identical to the per-rule entropy, except that the diff is taken from the baseline commit to the current checkpoint rather than from the previous checkpoint. So pi is file i's share of every line changed since the run began, and n is every file touched at least once.

In this harness. The same placement.change_entropy call, which computes both prefixes in one pass from two numstat invocations.

How to read it. An architecture where every rule lands in its own file keeps this high. An architecture where every rule lands in the same method keeps it low no matter how tidy any individual diff looked.

Share of all change so far in one file

cum_change_top1  ·  ↓ lower is better  ·  concentration ratio CR1 over the cumulative diff

Share of all change so far in one file

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Across every rule so far, what fraction of all changed lines landed in a single file. The most legible cumulative concentration number: this much of everything this project did happened in one file.

How it is calculated
CR1 = maxi linesi / Σj linesj, over diff(base, c)

Where. linesi is file i's cumulative changed-line count since the baseline commit.

In this harness. Same cumulative numstat as the cumulative entropy.

How to read it. This is the figure to quote when you want one sentence rather than an index.

Share of all change so far in five files

cum_change_top5  ·  ↓ lower is better  ·  concentration ratio CR5 over the cumulative diff

Share of all change so far in five files

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same, widened to five files, so a cosmetic split of the hot file cannot fix it.

How it is calculated
CR5 = Σi=1..5 lines(i) / Σj linesj

Where. lines(i) is the cumulative changed-line counts sorted descending, so the numerator is the five busiest files.

In this harness. Same cumulative numstat.

How to read it. If the top-1 share falls but this does not, the hot file was split rather than relieved.

Cumulative change concentration (HHI)

cum_change_hhi  ·  ↓ lower is better  ·  Herfindahl-Hirschman index over cumulative per-file changed-line shares

Cumulative change concentration (HHI)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The concentration index applied to history rather than to the current code. Two codebases can look structurally similar at the end while having got there very differently, and this is what tells them apart.

How it is calculated
HHI = Σi pi2,   pi = linesi / Σj linesj

Where. pi is file i's share of all lines changed since the baseline commit.

In this harness. placement.hhi over the cumulative numstat.

How to read it. The history view of concentration. It is the one metric here that a final-state snapshot cannot reproduce.

Careful. Not scale-free. An arm with more units scores lower for free, whatever the shape of its distribution. Always read it beside the Gini, which is scale-free, and beside the unit count. If Gini agrees, the honest claim is that the distribution is more unequal. If only this moves, the honest claim is only that the units are larger.

Files carrying the change so far

cum_change_files  ·  · descriptive  ·  harness, cumulative numstat

Files carrying the change so far

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How many distinct files have been touched at least once since the baseline. It is the denominator behind the cumulative concentration metrics, and a plain statement of how wide the project's footprint has grown.

How it is calculated
cum_change_files = | { i : linesi > 0 } | over diff(base, c)

Where. Counted over the cumulative numstat, so a file touched at change request 3 still counts at change request 60.

In this harness. Same cumulative numstat parse.

How to read it. Read it beside the cumulative HHI, which it deflates for free.

The structural-impact score

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Structural impact: the composite score

impact_composite  ·  ↓ lower is better  ·  defined in this harness (harness/metrics.py)

Structural impact: the composite score

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

One number for how much a change cost the structure. It is blast radius weighted by the complexity of the context that was disturbed. Editing a method inside a heavy god class costs far more here than the same edit inside an isolated unit, which is the whole design.

How it is calculated
cost(f) = max(WMC_other(f), 1) · CC(f) · max(1, Δlines(f))
impact_composite = files_changed · Σf ∈ changed cost(f)

Where. cost(f) is the per function term. WMC_other(f) is the summed cyclomatic complexity of the other methods in f's class, which stands for the context you must hold in your head to change f safely. A brand-new class has WMC_other = 0, which the max(…, 1) floors to 1, so a new isolated unit still costs something. That floor is what closes the fragmentation loophole. Δlines(f) is the changed-line count in f, and for a new function it is that function's own size. files_changed is the number of distinct production Java files the commit touched, applied as a multiplier, so scattering one rule across many classes is not free. The sum runs over every changed function, both modified and new. A within-commit rename is charged as a mutation rather than a free addition when the two bodies have a line-set Jaccard similarity of at least 0.6.

In this harness. metrics.impact_stats parses each touched file at both the previous and the current commit with lizard, matches functions by name, and falls back to body similarity for renames. Note what it cannot see: only lines inside a parsed function body count, so logic expressed declaratively, in a MapStruct expression, in openapi.yml, in schema.sql or in OfficeFloor's wiring, scores zero. Both arms have that escape hatch, so it is not an arm bias, but read a zero as “the logic went where this metric cannot look” rather than as “the change was cheap”.

How to read it. The scale is large and heavily skewed, because it is a product of four terms. The shape of the line matters far more than its value.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact: disturbing existing functions

impact_mutation  ·  ↓ lower is better  ·  this harness

Structural impact: disturbing existing functions

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The half of the impact score that comes from modifying code that already existed. This is the component that carries the discrimination between the two architectures, because the context weight means a mandated rule revision is genuine architectural signal rather than spurious re-touching.

How it is calculated
impact_mutation = files_changed · Σf ∈ modified ∪ renamed cost(f)

Where. cost(f) is the per function term. WMC_other(f) is the summed cyclomatic complexity of the other methods in f's class, which stands for the context you must hold in your head to change f safely. A brand-new class has WMC_other = 0, which the max(…, 1) floors to 1, so a new isolated unit still costs something. That floor is what closes the fragmentation loophole. Δlines(f) is the changed-line count in f, and for a new function it is that function's own size. files_changed is the number of distinct production Java files the commit touched, applied as a multiplier, so scattering one rule across many classes is not free. The sum is restricted to functions that existed at the previous commit and were modified, plus those detected as renames by the 0.6 Jaccard rule.

In this harness. Same single pass as the composite.

How to read it. This is where the two architectures separate most sharply. If the composite moves and this does not, the movement was all in additions.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact: adding new functions

impact_godclass  ·  ↓ lower is better  ·  this harness

Structural impact: adding new functions

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The half of the impact score that comes from new code: new files and new methods added to existing classes. It is reported so the composite's behaviour can be attributed to the right half.

How it is calculated
impact_godclass = files_changed · Σf ∈ new cost(f)

Where. cost(f) is the per function term. WMC_other(f) is the summed cyclomatic complexity of the other methods in f's class, which stands for the context you must hold in your head to change f safely. A brand-new class has WMC_other = 0, which the max(…, 1) floors to 1, so a new isolated unit still costs something. That floor is what closes the fragmentation loophole. Δlines(f) is the changed-line count in f, and for a new function it is that function's own size. files_changed is the number of distinct production Java files the commit touched, applied as a multiplier, so scattering one rule across many classes is not free. The sum covers functions with no counterpart at the previous commit. For these Δlines(f) is the new function's own line count, and WMC_other is 0 in a brand-new class, floored to 1.

In this harness. Same single pass as the composite.

How to read it. Because of the floor and the spread multiplier, both architectures pay something for additions. So this component does not separate them cleanly, and that is correct. The discrimination is supposed to live in the mutation term.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (composite): purely additive rules only

impact_composite_add  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (composite): purely additive rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that only add a rule and never revise an earlier one. This is the easy case, where an architecture is not asked to change its mind about anything.

How it is calculated
value(c) = impact_composite(c) if c is additive, else blank

Where. A change request is additive when the checkpoint plan declares no mutates list for it. The filtered field is left blank, not zero, on the other checkpoints. Blank matters: a zero would drag the mean of this view toward nothing on every revision checkpoint, which would make the two views incomparable.

In this harness. Derived in the gallery and in analyze from the base field and the checkpoint_type column. It is never written to the per-checkpoint CSV.

How to read it. If a condition only looks good here, it only looks good when nothing has to change.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (composite): rule-revision rules only

impact_composite_mut  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (composite): rule-revision rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that deliberately revise an earlier rule. Fourteen of the sixty change requests are of this kind, and they are placed deliberately rather than at random.

How it is calculated
value(c) = impact_composite(c) if c is mutative, else blank

Where. A change request is mutative when the checkpoint plan declares a mutates list, naming the earlier rules it is allowed to change. Additive checkpoints are blanked, for the same reason as above.

In this harness. Same derivation. Unlike the correctness metrics, the impact score keeps the mutative checkpoints in its base field rather than discounting them, because the context weight makes a mandated revision genuine architectural signal.

How to read it. The interesting case. Revising an existing rule is where an architecture either pays for having isolated the concern or does not. In the control run the concentrated arm paid roughly 36,000 per revision. The distributed arm paid 4,800.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (mutation of existing functions): purely additive rules only

impact_mutation_add  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (mutation of existing functions): purely additive rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that only add a rule and never revise an earlier one. This is the easy case, where an architecture is not asked to change its mind about anything.

How it is calculated
value(c) = impact_mutation(c) if c is additive, else blank

Where. A change request is additive when the checkpoint plan declares no mutates list for it. The filtered field is left blank, not zero, on the other checkpoints. Blank matters: a zero would drag the mean of this view toward nothing on every revision checkpoint, which would make the two views incomparable.

In this harness. Derived in the gallery and in analyze from the base field and the checkpoint_type column. It is never written to the per-checkpoint CSV.

How to read it. If a condition only looks good here, it only looks good when nothing has to change.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (mutation of existing functions): rule-revision rules only

impact_mutation_mut  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (mutation of existing functions): rule-revision rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that deliberately revise an earlier rule. Fourteen of the sixty change requests are of this kind, and they are placed deliberately rather than at random.

How it is calculated
value(c) = impact_mutation(c) if c is mutative, else blank

Where. A change request is mutative when the checkpoint plan declares a mutates list, naming the earlier rules it is allowed to change. Additive checkpoints are blanked, for the same reason as above.

In this harness. Same derivation. Unlike the correctness metrics, the impact score keeps the mutative checkpoints in its base field rather than discounting them, because the context weight makes a mandated revision genuine architectural signal.

How to read it. The interesting case. Revising an existing rule is where an architecture either pays for having isolated the concern or does not. In the control run the concentrated arm paid roughly 36,000 per revision. The distributed arm paid 4,800.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (new-function additions): purely additive rules only

impact_godclass_add  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (new-function additions): purely additive rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that only add a rule and never revise an earlier one. This is the easy case, where an architecture is not asked to change its mind about anything.

How it is calculated
value(c) = impact_godclass(c) if c is additive, else blank

Where. A change request is additive when the checkpoint plan declares no mutates list for it. The filtered field is left blank, not zero, on the other checkpoints. Blank matters: a zero would drag the mean of this view toward nothing on every revision checkpoint, which would make the two views incomparable.

In this harness. Derived in the gallery and in analyze from the base field and the checkpoint_type column. It is never written to the per-checkpoint CSV.

How to read it. If a condition only looks good here, it only looks good when nothing has to change.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Structural impact (new-function additions): rule-revision rules only

impact_godclass_mut  ·  ↓ lower is better  ·  this harness, filtered view

Structural impact (new-function additions): rule-revision rules only

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same impact score, restricted to the change requests that deliberately revise an earlier rule. Fourteen of the sixty change requests are of this kind, and they are placed deliberately rather than at random.

How it is calculated
value(c) = impact_godclass(c) if c is mutative, else blank

Where. A change request is mutative when the checkpoint plan declares a mutates list, naming the earlier rules it is allowed to change. Additive checkpoints are blanked, for the same reason as above.

In this harness. Same derivation. Unlike the correctness metrics, the impact score keeps the mutative checkpoints in its base field rather than discounting them, because the context weight makes a mandated revision genuine architectural signal.

How to read it. The interesting case. Revising an existing rule is where an architecture either pays for having isolated the concern or does not. In the control run the concentrated arm paid roughly 36,000 per revision. The distributed arm paid 4,800.

Careful. This score was defined on this experiment, so it cannot be the evidence for a claim about this experiment. It is also the quantity two of the four conditions were optimising, which makes it the clearest Goodhart demonstration here. Watch it collapse in those conditions while the published, externally defined metrics move far less.

Cohesion: does a class do one thing

Part of a series on how software architecture shapes AI driven code degradation. This post explains a group of measurements on their own. Each metric gets its definition, its figure, and its numbers.

Lack of cohesion, average class (LCOM)

ck_lcom_mean  ·  ↓ lower is better  ·  Chidamber & Kemerer 1994, via the CK tool

Lack of cohesion, average class (LCOM)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Counts the method pairs in a class that share no field, against the pairs that do. A class whose methods all touch the same state is one idea. A class whose methods touch disjoint state is several classes wearing one name. This is the numeric version of what the plain-English cohesion prompt asked for in words.

How it is calculated
LCOM(C) = max(0, |P| − |Q|)
ck_lcom_mean = mean over classes of LCOM(C)

Where. P is the set of method pairs in C whose accessed-field sets are disjoint. Q is the set of pairs that share at least one field. Low is cohesive. The measure is unbounded above and grows roughly with the square of the method count, so a large class is penalised twice: once for incoherence and once for being large.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure. The harness reads CK's per-class LCOM column and takes the mean, and separately the maximum.

How to read it. The natural place to check whether asking the agent for cohesion actually produced it.

Careful. Unbounded and method-count sensitive. Read it with the normalised LCOM* below and with tight class cohesion, which has the opposite sign. Agreement across all three is what makes a cohesion claim safe.

Lack of cohesion, worst class

ck_lcom_max  ·  ↓ lower is better  ·  C&K 1994, via CK

Lack of cohesion, worst class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The least cohesive class in the codebase. Unlike the mean it cannot be diluted by adding cohesive classes, so it answers whether there is a junk-drawer class in here, rather than whether classes are cohesive on average.

How it is calculated
ck_lcom_max = max over classes of LCOM(C)

Where. Same LCOM as above. Because it is unbounded and grows with method count, the worst class is often simply the largest one, so read it beside the class-weight concentration index.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. A step change here means a junk drawer appeared.

Lack of cohesion, normalised (LCOM*)

ck_lcom_star_mean  ·  ↓ lower is better  ·  Henderson-Sellers 1996, via CK

Lack of cohesion, normalised (LCOM*)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

A redesign of LCOM that is bounded roughly to the range 0 to 1 and does not simply grow with the number of methods. Because it is normalised, a difference here is a difference in shape rather than in class size, which is exactly the correction the original LCOM needs.

How it is calculated
LCOM* (C) = ( (1/a) Σj=1..a μ(aj) − m ) / ( 1 − m )

Where. m is the number of methods in C and a the number of fields. μ(aj) is how many of those methods access field j. So the first term is the average number of methods per field. If every method touches every field the numerator is m − m = 0 and the result is 0, meaning perfectly cohesive. If each field is touched by exactly one method the result approaches 1. Low is cohesive. Undefined for a class with fewer than two methods or no fields.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. The version to prefer when comparing arms whose classes differ in size.

Tight class cohesion (TCC)

ck_tcc_mean  ·  ↑ higher is better  ·  Bieman & Kang 1995, via CK

Tight class cohesion (TCC)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The fraction of method pairs in a class that are directly connected through shared field access. It is reported precisely because its direction is inverted relative to LCOM: if a condition improves LCOM and worsens TCC, the improvement is an artefact of one definition rather than a real gain in cohesion.

How it is calculated
TCC(C) = NDC / NP,   NP = m(m − 1) / 2

Where. m is the number of visible methods. NP is therefore every possible pair of them. NDC is the number of pairs that are directly connected, meaning they access at least one instance variable in common. High is cohesive, which is the opposite sign to LCOM. Undefined, and blank, for a class with fewer than two methods.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Agreement with LCOM in the opposite direction is what makes a cohesion claim safe. Disagreement means one definition is doing the work.

Loose class cohesion (LCC)

ck_lcc_mean  ·  ↑ higher is better  ·  Bieman & Kang 1995, via CK

Loose class cohesion (LCC)

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The same as tight cohesion, but it also counts methods connected indirectly, through a chain of other methods. It is always at least as high as the tight version, and the gap between them is informative on its own.

How it is calculated
LCC(C) = (NDC + NIC) / NP

Where. NIC is the number of pairs connected only indirectly: not sharing a field themselves, but linked through a chain of methods that do. NDC and NP are as in TCC. High is cohesive.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. A large gap between LCC and TCC means the class holds together only through intermediaries, which is weaker cohesion than the LCC figure alone suggests.

Lack of cohesion of the handler class

ck_handler_lcom  ·  ↓ lower is better  ·  C&K 1994, via CK, pinned to the handler class

Lack of cohesion of the handler class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

Cohesion of the one class both architectures agree is the entry point. Codebase averages can be moved by adding files. This cannot.

How it is calculated
ck_handler_lcom = LCOM(H)

Where. H is the handler class, matched by name in CK's per-class output. Same LCOM definition as the codebase mean. Low is cohesive.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. The like-for-like cohesion comparison. A controller accumulating rules that touch disjoint state climbs here.

Tight cohesion of the handler class

ck_handler_tcc  ·  ↑ higher is better  ·  Bieman & Kang 1995, via CK, pinned to the handler class

Tight cohesion of the handler class

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

The inverted-sign cohesion check, on the handler class.

How it is calculated
ck_handler_tcc = TCC(H) = NDC(H) / NP(H)

Where. Same TCC definition. High is cohesive. Undefined, and therefore blank, for a handler class with fewer than two methods, because there are no pairs to connect.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Sparse by construction in whichever arm keeps its handler minimal. The blanks are a result, not missing data: a handler with one method has no cohesion to measure.

Depth of inheritance tree

ck_dit_mean  ·  ↓ lower is better  ·  Chidamber & Kemerer 1994, via CK

Depth of inheritance tree

Line is the mean of ten runs. Band is one standard deviation. Click for full size.

How deep the class hierarchy goes, on average. It is reported to show that neither architecture is achieving its structure through inheritance, which matters because none of the other metrics on this page would attribute complexity hidden in a hierarchy correctly.

How it is calculated
DIT(C) = number of edges from C up to the root
ck_dit_mean = mean over classes of DIT(C)

Where. A class extending nothing has DIT = 1 in CK's convention, counting Object as the root. Interfaces implemented do not add depth.

In this harness. Computed by the CK tool, which parses source with Eclipse JDT. Only declared source classes are seen: library types are not, and neither is anything the parser fails on, so the class count in the validity group is worth checking beside any CK figure.

How to read it. Flat and low in both arms is the expected and desired result. A rise would mean complexity moved into a hierarchy.