Insights11 min read

What is spec-driven testing? Two readings of one spec

By qtrl Team · Engineering

Spec-driven development tools all follow roughly the same sequence. You write a specification, the agent turns it into a plan, the plan becomes a task list, and the tasks become code. GitHub's launch post for Spec Kit is clear about where the human fits in: "your role isn't just to steer. It's to verify."

Verify how, though? The spec gets a lot of tooling. The plan gets a lot of tooling. The verification step mostly gets you, reading a diff and a set of tests the same agent wrote ten minutes earlier.

Spec-driven testing is the name for doing that last step properly. We've already written about what spec-driven development changes for QA. This post is about the testing half: what it is, where it came from, and why we think end-to-end tests are the strongest form of it.

A definition you can use

Spec-driven testing means deriving, running and maintaining tests from a written, reviewed specification of intended behavior. The spec, not the implementation, decides what the right answer is. Every test points back to the part of the spec it checks, and every result can be traced to that same line.

The term isn't settled yet. Some vendors use it as a new label for writing tests in plain English. A 2026 Google paper uses it for an agent that writes down a function's pre-conditions and post-conditions before generating unit tests. And testing standards have used "specification-based testing" for decades to mean black-box techniques like boundary values, equivalence partitions and decision tables, all derived from the functional spec.

Those meanings overlap more than they compete. The classic one tells you how to pick test cases from a spec. The new one is about who turns the spec into running tests, and how those tests stay tied to the spec as both change.

One spec, two independent readings

When a team builds from a spec, the spec gets read twice. Once by whoever writes the code, who has to decide how the system will behave. Once by whoever writes the tests, who has to decide what must be true when it's done. If those two readings happen independently and then agree on the running product, you have something much stronger than either one alone. You have two separate interpretations of the same requirement that landed in the same place.

If they disagree, that's useful too. Either the code misread the spec, the test misread the spec, or the spec was ambiguous enough to support both readings. Kiro builds this exact choice into its property-based testing: when a property derived from the spec fails, it asks whether to fix the code, the spec or the property. All three outcomes improve quality. The worst case is the one where nobody finds out.

One spec, two independent readingsThe reviewed specacceptance criteria, user-visible behaviorreading 1reading 2Implementationcoding agent or developerdecides how it worksEnd-to-end testsseparate author, never sees codedecides what must be trueno sharedcontextdeployedexercised in a real browserThe running applicationthe only place both readings meetThey agreeindependent evidence the buildfollows the specThey disagreethe code, the test, or the specis wrong, and now you know

The word doing the work in all of this is independent. Our earlier post called out the trap where code and tests come from the same document and confirm each other's mistakes. Deriving both from the spec is fine. Deriving them in the same head, human or model, is where it goes wrong.

Independence doesn't happen by accident, either. In 1986, John Knight and Nancy Leveson had separate teams write versions of the same program from one specification and found the versions failed together far more often than chance would predict. People make similar mistakes on the same hard parts. LLMs are worse here, because the same model tends to make the same mistake every time. A 2025 study found that test suites written by an LLM reflect the generating model's own error patterns.

So if one agent, in one session, reads the spec, writes the code, and then writes the tests, you don't have two readings. You have one reading, copied. For the second reading to be independent, it needs:

  • A different author or context. A separate agent session, a separate tool, or a person. Not the one that just wrote the code
  • No access to the implementation. The moment the test writer reads the code, it starts asserting what the code does instead of what the spec says
  • Expected results nobody downstream can edit. If an agent can rewrite an assertion to get to green, the test stops being a check
  • A spec a human actually reviewed. Both readings inherit whatever the spec gets wrong, so the spec review is where the shared risk gets caught

The second point has real evidence behind it. Across 24 Java projects, Konstantinou et al. found LLMs were more likely to write test oracles that capture what the program actually does than what it should do. Bugs included. A small 2026 pilot that rewrote ten known bugs as business requirements found the opposite pattern when the model only saw the requirement: its expectations agreed more with the intended behavior than with the buggy code. Ten bugs is a pilot, not proof, but the direction is consistent.

Why end-to-end tests are the strongest independent check

Specs are written in the language of the user. "When a signed-in user submits an expired coupon, the checkout shows an error and the total doesn't change." Nothing in that sentence mentions a function, a table or a service.

Unit tests can't check that sentence directly. They check pieces of the implementation, which means they're written in the code's vocabulary and shaped by the code's structure. Someone had to decide that coupon validation lives in this function, returns this type, and gets called from there. That decision belongs to the implementer's reading of the spec. You still want unit tests for speed and precision. They just sit inside the first reading. They verify that the parts work the way the builder intended.

An end-to-end test sits outside it. It observes the same surface the spec describes: what the user does, and what the user sees. It runs through the real stack, with the real configuration, routing, integrations and data. It doesn't care how the code is organized, so a refactor doesn't change it, and it can be written before any code exists. It's the only kind of test that can be derived from the spec alone and still check the whole claim.

That's why we'd call end-to-end tests the closest thing testing has to independent proof that an implementation follows its spec. It can't tell you the spec was right (more on that below). What it can tell you is that what got built matches what got asked for, checked by something that never saw how it was built.

The historical objection to leaning on E2E tests was cost. They were slow to write, slow to run, and fragile. The writing part has changed a lot in the last two years, which is most of why this topic is coming up now.

This has been tried before

Making the spec executable is an old idea, and its history explains a lot about what can go wrong.

Ward Cunningham's Fit (2002) let customers write examples in tables that programmers wired up to the software. BDD grew out of TDD in the mid-2000s, and Cucumber (2008) gave it Given/When/Then scenarios that read like requirements and ran like tests. Gojko Adzic's Specification by Example (2011) put the emphasis where it belonged: the conversation that produces the examples, with regression tests as a by-product.

The people who built these tools wrote the sharpest criticism of how they got used. Aslak Hellesøy, who created Cucumber, wrote in 2014: "If you think Cucumber is a testing tool, please read on, because you are wrong." Teams wrote scenarios after the code, on their own, and ended up with slow and brittle suites nobody outside engineering ever read.

The more rigorous branches, model-based testing and property-based testing, worked well technically and stalled on authoring cost. Writing a good state model or a good property is specialist work. Contract testing, with Pact or an OpenAPI schema, is probably the most successful spec-driven testing in production today, and it's telling why: the spec is narrow, machine-readable, and small enough that people actually keep it current. We cover that in how to do contract testing.

So two things held these practices back. Authoring the executable spec was expensive, and nobody read the spec once it existed. LLMs mostly fix the first problem. They do nothing for the second, and by producing more text faster they can make it worse.

What the SDD tools do about testing today

The spec-driven development tools take testing seriously, and they deserve credit for it. Spec Kit's default constitution makes test-first development non-negotiable and turns acceptance scenarios into tests. Kiro writes requirements in EARS notation and, since its general availability release in November 2025, can generate property-based tests that link back to requirement numbers. OpenSpec validates spec structure. Tessl tags spec items with the tests that cover them.

Look at where verification lands across all of them, though. It falls into three buckets: checking the spec itself for gaps and contradictions, tests the coding agent writes from the spec, and human review. What's mostly outside their scope is a separate, independent check of the running application against each acceptance criterion. That's not a flaw in those tools. They're development tools, built around the first reading. The second reading is a different job.

How an agent runs a test written in English

Two designs dominate, and they end up in a similar place.

The first generates code once and repairs it when it breaks. Playwright's test agents work this way: a planner explores the app and writes a Markdown test plan, a generator turns it into test files, and a healer fixes failures or marks a test as skipped when the feature itself looks broken. We wrote up how those agents fit together.

The second keeps the natural-language test as the source and has an agent execute it against the browser, then records or caches the actions so later runs replay without calling the model on every step. The model comes back in when a step no longer matches the page.

Either way, the model does its thinking while a test is being written or repaired, and the routine run is as deterministic as possible. That matters for spec-driven testing, because a verdict that changes between identical runs isn't evidence of anything. Self-healing is part of the same picture. It's good at absorbing changes that don't affect meaning, like a renamed button or a moved field. A heal that changes what the test expects is a different thing, and it should go to a human as a question about the spec. Our self-healing explainer covers where that line sits.

What the evidence says so far

The research is early, and most of it sits below the browser.

The strongest result is the Google paper mentioned above. Having an agent write a contract-style spec before writing tests found 9.8 percentage points more bugs (p = 0.035) and added 2.5 points of branch coverage on Google production bugs, compared with an agent that went straight to writing tests. Those were unit tests on one company's code, but the mechanism (reason about what should be true, then test it) is exactly the one spec-driven testing relies on.

Meta's experience with TestGen-LLM shows the other half. It only kept tests that built, passed reliably and added coverage, and engineers accepted 73% of its recommendations. The filters mattered as much as the model.

For end-to-end tests generated from natural-language specs, there isn't yet a large study of bug detection, flake rate or cost. We haven't found one, and we'd be suspicious of anyone quoting a precise number. What the evidence does show is a consistent direction: expectations derived from intent catch more than expectations derived from code.

Where it goes wrong

The spec is wrong. Both readings inherit it. A spec that leaves out a case produces code that leaves it out and tests that never ask about it. Spec-driven testing proves conformance, and only exploratory testing and a good spec review catch a spec that describes the wrong thing.

The spec is too long to review. Birgitta Böckeler, reviewing SDD tools on martinfowler.com, put it bluntly: "I'd rather review code than all these markdown files." A spec nobody reads gives you no independent assurance at all. Acceptance criteria are the right size for testing. Design documents usually aren't.

The model is asked to find the ambiguity. LLMs can flag unclear requirements, but a 2026 study measured precision between 41% and 61% on that task. Useful as a list of questions for a person. Not useful as a gate.

The agent edits the oracle. Kent Beck has described coding agents that delete or disable failing tests to get to green. Whatever implements the spec shouldn't be able to touch the expected results that check it. Our guide to reviewing AI-generated tests has the patterns to watch for.

The spec drifts. This is where spec-driven testing flips from risk to remedy. One review of OpenSpec sums up a gap most SDD workflows share: "nothing keeps specs synced with code." A suite derived from the spec and run on every change is the drift detector. When it fails, someone has to decide which side moved.

Where this is heading

For twenty years, most test oracles came from one of two places. The implementation (record what it does and assert that) or expensive formal artifacts like models, properties and contracts that few teams had time to write. LLMs make a third source cheap: the plain-language spec the team already agreed on.

That turns the spec into the test plan, not just the build plan. The teams that get the most out of spec-driven development will be the ones that treat it as two jobs. One reading builds the product. A second, independent reading checks it in the browser, where the user will check it anyway.

Next up: how to put this into practice, from acceptance criteria to browser evidence, in how to turn acceptance criteria into E2E tests.

Have more questions about AI testing and QA? Check out our FAQ