How-To7 min read

How to turn acceptance criteria into E2E tests

By qtrl Team · Engineering

Most user stories already contain a test plan. It's the acceptance criteria section, the part that says what has to be true before anyone calls the story done. In a lot of teams that section gets read once during refinement, once by whoever builds the feature, and then never again.

In part one we argued that spec-driven work gets its quality from two independent readings of the same requirement. One reading builds the feature. The other checks it, from the outside, in a real browser. When both land in the same place, you have evidence the implementation follows the spec. When they don't, you've found a bug in the code, the test or the spec.

This post is the practical half. It walks through turning acceptance criteria into end-to-end tests that stay independent from the code, using qtrl as the example. The principles carry over to whatever tooling you use.

From acceptance criteria to browser evidenceRequirementstory or ticketwith its criteriaDraft testgenerated fromthe requirementFirst runsteps adapt,expectations lockedReviewdraft, pending,approvedEvery runstrict, againsteach environmentDeviation logevery changed step, with a reasonRequirement IDs, versions and audit logevery run, screenshot, video and report traces back to the criterion it checks

Step 1: Write criteria a browser can check

An end-to-end test can only check what a user can observe, so the criteria have to be written at that level. "The coupon service validates expiry" is an implementation note. "When a user applies an expired coupon at checkout, an error explains the coupon has expired and the order total doesn't change" is something a browser can verify.

A few habits help:

  • One behavior per criterion. If a criterion has "and" joining two outcomes that can fail separately, split it
  • Name the trigger and the visible result. The EARS patterns from our spec-driven development post (WHEN this happens, THE system SHALL do that) are a good template
  • Include the unhappy paths you care about. Generated tests follow the criteria they're given. If an expired token, an empty cart or a missing permission matters, say so
  • Give each requirement an ID. A Jira key works. You need something stable to trace back to

Do this in the spec review, before anyone builds anything. That's the one place where both readings can be protected from the same mistake. For more on the craft, see how to write test cases from requirements.

Step 2: Generate the test from the requirement, not the code

Independence starts with what the test author is allowed to see. In qtrl, test generation starts from the requirement text. You can paste a requirement into the app, send it through the API, or start from a Jira issue, where qtrl builds the input from the summary, description and acceptance criteria fields.

The generator combines that requirement with what qtrl has learned about your application from earlier runs and exploration: which pages exist, how to get between them, what the forms look like. What it doesn't see is your source code. qtrl tests the running application from the outside, so the tests it writes are a reading of the requirement, not a reading of the implementation. That's the separation part one was about, built into how the tool works.

Generate one story at a time. A focused requirement with three to six clear criteria produces a better test than a whole epic pasted in at once, and long inputs get trimmed anyway (the Jira input is capped at 2,000 characters).

Step 3: Let the first run fix the steps, never the expectations

A test written from a requirement is a draft. It knows what should happen but has to guess some of the clicks: the exact label on a button, whether the coupon field sits behind a toggle. So the first execution of a new test in qtrl runs in adaptive mode against the real application.

In that run, the agent can change the steps to match the actual UI. It can insert a step, drop one, or reword one. It is explicitly not allowed to change the expected results, because those came from the requirement and are the whole point of the test. Every change it does make is recorded as a deviation: which step, what it said before, what it says now, and why.

Read the deviations. Most will be harmless, like "the button says Apply, not Submit." Some will be the two readings disagreeing. If the agent had to work around a missing confirmation message, or couldn't find a field the spec assumes exists, that's a question for the team, not a detail to wave through. It's the same three-way call from part one: fix the code, fix the spec, or fix the test.

Step 4: Review it before it counts

A generated test shouldn't join your regression suite just because it ran. In qtrl, tests move through Draft, Pending Review and Approved (or Rejected). After the first run, an AI review step checks whether the test still covers what it set out to cover, so a draft that drifted into testing something else gets flagged instead of quietly approved.

Then a person looks at it. The questions are short: does each expected result match a criterion, are the unhappy paths there, and do the deviations make sense? Treat it like a pull request. Our guide to reviewing AI-generated tests covers the patterns that slip through, and most of them apply here.

Step 5: Keep the thread back to the requirement

Add the requirement ID to the test. qtrl stores requirement IDs on each test case, so a test for PROJ-412 says so, and anyone looking at a failure can go straight to the criterion it checks.

Tests also keep their history. Every edit creates a new version you can restore, and changes land in an audit log. When a requirement changes, that history is how you show which version of the test checked which version of the spec. If you work somewhere that needs requirement-to-test traceability for audits, this is the part that saves a spreadsheet.

Step 6: Run it on every environment and every change

Once a test is approved, runs switch to strict mode. The agent follows the steps as written and checks each expected result. If a step can't be executed as written, it fails the step and says why. It doesn't improvise its way to a pass. That's what you want from a check that's supposed to be independent: when the product stops matching the spec, the test goes red.

Test runs in qtrl target a specific environment (development, test, staging or production), and each execution can go to the AI agent or be assigned to a person. To make this part of your pipeline, the REST API lets you create a run, execute it and fetch the report with a project-scoped API key. Running the spec-derived suite on every deploy to staging is what turns it into a drift detector. Our post on agentic testing in CI covers where those runs fit in a pipeline.

Step 7: Read the evidence, then decide which side moved

Every run records per-step results, screenshots and a video, and you can export a PDF report of the whole run. That evidence is what makes an end-to-end result useful beyond pass or fail. You can see what the user would have seen.

When a spec-derived test fails, don't edit the test until it passes. Look at the screenshot, read the criterion, and decide which of the three it is:

  • The code is wrong. The build doesn't do what the spec says. File the bug against the requirement ID
  • The spec changed, or was wrong. Update the requirement first, then the test, so both readings move together and the history shows why
  • The test is wrong. It misread the criterion or checks something too narrowly. Fix it and send it back through review

Only the last case should end with you editing an expected result. If that's happening often, look at how the criteria are written in step 1.

What stays outside this loop

Spec-driven tests prove the product matches the spec. They can't tell you the spec described the right product. A missing requirement produces a missing test, and the suite stays green. That's what exploratory testing is for, and it matters more, not less, once most of the scripted checking is generated.

End-to-end tests also don't replace unit tests. Developers still need fast, precise feedback on the parts they're building, and that belongs close to the code. The spec-derived E2E suite answers a different question, the one the business actually asked: did we build what we said we'd build?


qtrl generates browser tests from your requirements and Jira acceptance criteria, adapts the steps to your real UI without touching the expectations, and keeps every run traceable to the requirement it checks, with screenshots, video and a full audit trail. Your code stays one reading of the spec. qtrl gives you the second one. Try it on your next story.

Have more questions about AI testing and QA? Check out our FAQ