The paper is out as a preprint. Version two follows when the review arm finishes its run. That work answers a question about architecture. It does not answer the question that decides whether any of this is usable, which is whether the whole loop can run without a developer sitting in it.
So here is the loop with a column for where each part of it stands.
| step | what it does | where it stands |
|---|---|---|
| Clarify | Turn what the user said into what the user meant, by asking them. | I have not built it. Others are working on it. |
| Specify | A refined change request, in English, that one person could hand to another. | Held fixed in all my runs, as the thing being varied. |
| Test | Write the acceptance test that pins the requested behaviour, before any code exists. | New repository. Instrument built, nothing graded yet. |
| Implement | Make the test pass without breaking the suite. | Two harnesses, hundreds of runs, one preprint. |
| Review | An independent reader critiques the change; the author acts on it. | Arm built, run in progress. |
The gap in the middle of that table is what the rest of this is about.
What I have measured
The REST harness identified the amount of complexity is conserved and the placement of it is not. A mutative architecture ends up wherever the last prompt, tool or metric pushed it. An additive one stays put, because there is no slack for an instruction to take up.
The obvious objection is that this is one endpoint in one language. The full stack harness covers the other half. It starts from a near empty base (a front end shell, Spring with the OfficeFloor plugin, an in memory database with no tables) and grows a whole application through about sixty plain English change requests, each one spanning a schema migration, the server, the front end and its acceptance test. Erosion is tracked per layer, TypeScript as well as Java, and the finished chains run on OfficeFloor + TanStack and OfficeFloor + React, with Angular and HTMX stacks built to the same contract. The claim is about composing a whole application, so a result that held only for the server would not be worth much.
One decision in that harness carries everything else. The tests bind only to data-testid attributes and assert only through the running UI with Playwright. Data is arranged through a profile guarded seed endpoint, so a test arranges but never asserts through a back door. The whole stack underneath (schema, server, components) can be regenerated and the test still means what it meant before. Functionality is tracked as tests rather than as specifications, and sixty changes later the suite is the only artefact that survived untouched.
Which makes the suite the weak point
In both harnesses I wrote the tests. That is the right call for an experiment about architecture, since you do not let the thing under test grade itself. However, it quietly assumes away the step the pipeline depends on. In the real pipeline nobody hands the agent a known good acceptance test. The agent writes it before the feature exists.
And a wrong test is worse than no test. A missing test leaves a gap you can see. A test that passes without asserting the thing it was written for turns the gate green and lets the pipeline proceed. Every later change is then protected by something that is not protecting anything. Over three to five hundred unattended changes that is the failure that compounds.
So the new repository holds the same sixty checkpoints with the halves swapped: the application fixed and known correct and the agent writing the tests.
This is not new ground and that is useful
Measuring generated tests by breaking the program and seeing whether a test notices is the standard approach, and it has been done to LLM test generation plenty of times. Meta's mutation guided work puts a plain LLM prompt at roughly 53% mutation score on a Java benchmark and reaches about 89% by feeding the surviving mutants back in. There is a study specifically on coding agents writing over mocked tests, which is my failure mode arrived at from another direction. Deriving the mutations from requirements rather than from syntax is not mine either, since specification mutation dates to the late nineties.
Good. It means there are numbers to compare against instead of a blank page. What I could not find measured together is the particular shape I need:
- Acceptance tests through a running UI, rather than unit tests on benchmark functions. Mutating a whole full stack application and grading it through a browser is rare, because it is expensive: every mutation needs a build, a serve and a Playwright run.
- Written before the implementation exists. Most generated test work runs against code that is already there, so the test can be checked by running it. Here the agent sees the previous checkpoint and never the current one. It cannot confirm its new test passes. That asymmetry is the hard part of test first work and the main source of bad tests.
- Sixty consecutive changes to one application. The long horizon benchmarks ask how the code degrades. This asks whether the suite stays honest and what it costs to keep.
It found four holes before any agent ran
Calibrating the instrument means running my known good suite against every mutation and checking each one dies. Four did not.
Four checkpoints asserted less than their request promised. The sharpest was checkpoint 21, where the request was "for each one give a description, how many, and the price each. Work out the total for me", and the spec asserted the line description and the invoice total but never a line's own quantity, unit price or amount. A line amount computed without its quantity, which is correct whenever the quantity happens to be one, was invisible to the suite that existed. Three others were the same shape.
Those were tests written with the requests available. The hole rate on careful human work is not zero and that is the bar the agent is measured against rather than perfection.
No comments:
Post a Comment