Figr is the AI product designer that understands your product.
Try for freeSee a demo
Guide

Test Case Generation from Requirements That Ships

Test Case Generation from Requirements That Ships
Published
October 9, 2026

The first mistake in test case generation from requirements is treating it like a prompt problem when it is really a requirements problem.

A 2025 survey of 267 studies found that only 39% of requirements-based test generation approaches created test cases directly from requirements, while 61% first transformed the requirement into a model, formal specification, or structured representation (survey of RBTG studies). That split highlights something many teams miss: if the requirement is fuzzy, no amount of automation will rescue it, and a polished test can still hide an untestable product decision.

The useful move is earlier and sharper. Treat generation as a diagnostic for requirement quality, then use the resulting structure to produce tests you can trace, review, and trust.

Why Most Test Generation Fails Before Generation Starts

A checkout story can look complete and still fail at the point that matters. A product team I watched had an acceptance criterion for applying a discount code, but nobody had defined what should happen when the code expired, the cart changed, or the session dropped. Test generation could not repair those gaps. It only made them more visible.

Practical rule: if the requirement can't tell you what should fail, it can't yet tell you what should pass.

Requirements defects are expensive because they are expensive to correct late. The Software Engineering Institute summarizes findings that these defects can cost 10 to 200 times more to fix after release than during requirements development (SEI on requirements defects).

The cost signal hidden inside every “simple” request

That cost profile changes how test generation should be used. An underspecified requirement does not become clearer because a model produced more tests. It usually becomes faster to misread.

Review rooms make this obvious. Teams celebrate velocity because an AI drafted twenty cases, but the cases only mirror the sentence they were given. If the sentence never defined failure behavior, the test set cannot reliably separate a business rule from a guess.

The warning shows up in requirements inspections too. NASA's Jet Propulsion Laboratory found higher defect density in requirements inspections than in later lifecycle products, with correctness, logic, and completeness appearing often in those documents. That is a diagnostic signal, not just a quality statistic. Weak requirements create shallow tests because the generator has no stable decisions to work from.

The better question is whether the requirement is specific enough to test at all. If it is not, the first output should be clarification questions and defect notes, not test cases. That shift turns test generation into a requirements-quality check before it becomes a production workflow.

Exit criteria help keep that check concrete: zero requirements without defined failure behavior, and every acceptance criterion should include explicit invalid and expired states before generation starts. For teams trying to tighten product language before QA ever sees it, clarity for product teams can help establish the kind of spec discipline that makes testing possible.

Parsing and Normalizing Requirements into a Test-Ready Model

A usable workflow starts by turning rough product language into a structured representation. A survey of RBTG studies found that 61% of approaches first transformed requirements into a model, formal specification, or other intermediate form, which is why the model layer is where the key decisions happen. Once you normalize the requirement, the test stops being a guess and becomes a traceable artifact.

A six-step path from sentence to test

1. Identify the moving parts.
Pull out actors, preconditions, triggers, states, decisions, and expected outcomes.
If those pieces are not visible in the requirement, the requirement is not test-ready.

2. Assign a stable requirement ID.
One ID, one source of truth.
That reference keeps QA, product, and engineering aligned on the same object.

3. Classify ambiguity before you write a test.
Look for vague terms, missing conditions, conflicting statements, and underspecified actors.
These are the signals that tell you whether clarification is needed.

4. Model the paths.
Write down normal, alternate, exception, and boundary flows.
That exposes what the user can do, and what the system should refuse.

5. Generate abstract cases.
At this stage, the output should describe behavior, not data values.
The test purpose should be clear before inputs are finalized.

6. Concretize and trace.
Add data, environment, and oracle assertions, then link each case back to the requirement ID.
A generated test that cannot be traced is just a note with better formatting.

A test case is useful only after it can survive review, execution, and change.

A concrete example makes the sequence easier to see. “Users can update their shipping address” is not enough. A test-ready version has to define who the user is, when the edit is allowed, what happens if the address is incomplete, and whether a saved order can still be modified. Once that structure exists, case generation becomes controlled translation, not creative interpretation.

For teams that want examples of how requirement artifacts are usually written, browse PRD and SRS templates before you lock the model. If the source document is still being shaped, writing a product requirements document with testable structure in mind reduces churn later.

A checklist infographic titled Discovering Edge Cases that outlines ten critical scenarios to consider in product design.

Discovering Edge Cases That Slip Past Product Reviews

Teams usually miss edge cases because they validate the happy path first and treat unusual states as follow-up work. A better sequence is to force each requirement through the failure states that recur across products, then decide whether the requirement is ready for design sign-off. Figr's edge-case framing fits that approach because it treats review as a requirements-quality check, not just a documentation step.

The ten categories worth checking every time

Empty states. What should the user see when there's no data yet?
Miss this, and the first live screen looks broken.

Loading sequences. What happens while the system is still working?
A spinner without context often becomes a support ticket.

Boundary values. What happens just below and just above the limit?
If a field accepts a quantity, the off-by-one case is where bugs hide.

Invalid inputs. What if the input is malformed or out of range? This is when validation becomes visible.

Permission failures. What if access is blocked?
The user needs a clear refusal, not a silent dead end.

Network errors. What if the connection drops or slows down?
Recovery matters as much as success.

Concurrent edits. What if two people change the same thing at once?
Product policy shows up in the UI.

Timeout paths. What if the operation takes too long?
The user needs a state change, not indefinite waiting.

Cancellation. What if the user aborts midway?
The system should know how to stop cleanly.

Recovery. What if the flow has to restart after failure?
Trust is either rebuilt or lost here.

Run each acceptance criterion through those ten states and record what is missing. Exit criteria: each acceptance criterion has been walked through all ten states and gaps are logged as findings with requirement IDs before design sign-off. If a state does not exist in the design, that absence is a finding, not a small note to revisit later. If a flow has to go public before every gap is closed, a public roadmap can help teams expose what's planned without pretending the missing cases do not matter.

The payoff is usually concentrated in the same places. Edge cases cluster around permissions, interruption, and recovery, so design review is the cheapest point to catch them. A case found there is a short correction cycle. A case found after launch becomes a support burden, a defect, and a product trust problem at the same time.

Generating Executable Test Cases from Designs and PRDs

A generated suite is only as good as the requirements behind it. Teams that start from a thin brief often get polished test wording and weak coverage, because the missing detail was already in the source. For teams formalizing the source document before generation, see writing a product requirements document as a companion reference.

Screenshot from https://figr.design

A practical workflow starts with the live product context. A PM reviews a checkout flow, captures the app in Chrome, and gives the system the actual screens, states, and transitions instead of a detached summary. From there, the generated set can include acceptance tests, failure cases, and QA handoff artifacts that stay tied to the requirement they came from.

What a useful generated case actually contains

It should not stop at “log in and click checkout.” The case needs the preconditions, steps, expected result, and the requirement or acceptance-criterion ID that justifies it. If those fields are missing, QA has to reconstruct the intent later, which makes review slower and traceability weaker.

The better output also reflects the product context around the flow. If checkout depends on a discount code field, the generated set should include empty input, invalid codes, and recovery after error. If the screen uses design-system tokens or raises accessibility concerns, those should appear in the same artifact, because a separate review often catches them too late.

Figr is useful here because it generates PRDs, user flows, edge cases, and test cases from live product context, then keeps the output grounded in the app rather than in a blank template. That matters when a team needs a first-pass QA suite that reflects what the interface does, not what the prompt hoped it would do.

For teams formalizing the source document before generation, developer handoff best practices are a good companion reference because the handoff quality often decides whether the test suite is usable or just complete on paper.

QA Handoff Templates and Traceability That Survive Review

A generated case becomes operational only when QA can run it, trace it, and defend it. That means the handoff package needs enough structure to survive a real review, not just an export. The cleanest way to do that is to keep the requirement ID visible at every step.

The seven fields that matter

Requirement ID. The anchor for change tracking and traceability.
Without it, the test floats free from the business need.

Preconditions. What must already be true before execution.
This prevents false failures caused by setup drift.

Steps. The exact actions the tester or automation should take.
If a step is vague, the case is not executable.

Expected result. The observable outcome the team will judge against.
This is the oracle in plain language.

Oracle. The rule or evidence source that defines pass or fail.
It keeps “looks right” from replacing verification.

Priority. Where the case sits in the release queue.
Not all cases deserve the same urgency.

Risk tag. The reason this case matters now.
Product judgment enters the test list here.

The ISTQB syllabus says acceptance criteria should be verified by acceptance tests and that traceability should run between the requirement or user story and the related test cases (ISTQB Certified Tester Specialist syllabus). NIST adds a conformance-oriented structure, where you define assertions, derive a test purpose for each one, and then create the test case that executes that purpose (NIST conformance-oriented process). Put those together, and the handoff becomes a chain: requirement, assertion, purpose, case, execution result.

That chain matters for another reason. Requirements coverage and statement coverage answer different questions, so they shouldn't be collapsed into one vague score. For teams that report execution coverage against requirements, traceability is what makes the report credible (ISTQB sample answers on requirements coverage). The practical rule is to keep the requirement reference in the case itself, then update execution status as the suite runs.

If a test can't point back to a requirement, it can't help you spot a missing requirement.

Measuring Coverage Without Fooling Yourself

Coverage numbers are easy to admire and easy to misread. The trap is treating a higher percentage as proof of better fault detection. In a NASA study spanning four industrial systems, researchers evaluated 15,000 combinations of test suites, case examples, and mutant sets, and found no statistically supported hypothesis that the coverage criteria they examined, by themselves, guaranteed better fault finding (NASA study on coverage and fault finding).

An infographic titled Measuring Coverage Without Fooling Yourself detailing best practices for accurate data coverage metrics.

Why one number is never enough

NASA experimentation also showed how requirements coverage can overstate assurance. One case achieved 65.71% requirements coverage, but only 43.65% statement coverage and 20.01% precise checked coverage. The gap is the point. A requirement can be marked covered while the executable behavior is still only partly exercised, so the report looks stronger than the suite really is.

A better coverage readout combines several signals:

  • Requirements coverage shows whether the behavior is represented.

  • Branch or statement coverage shows how much implementation ran.

  • Mutation score shows whether the tests detect seeded faults.

  • Oracle quality shows whether the pass or fail decision is meaningful.

  • Escaped-defect yield shows whether the suite catches the failures that matter.

Review the failure modes directly as part of the same check. Ambiguous natural language, missing exception paths, non-deterministic expected results, unbounded combinatorial input spaces, and abstract tests that cannot be mapped to executable steps all weaken the suite before it runs. A useful review asks whether the case is covered, observable, stable, and still worth keeping. For version-aware governance, report requirements coverage, branch coverage, and mutation score together for each release version, and do not let a release pass on requirements coverage alone.

One infographic, Measuring Coverage Without Fooling Yourself, captures the practical standard: coverage metrics are useful only when they are interpreted together and tied back to the requirement model.

In short, one coverage number encourages theater. A small set of linked signals supports judgment.

Keeping Tests Alive When Requirements Change Every Week

The hardest part of test case generation from requirements is not the first draft. It's keeping the suite honest after the requirement changes. Continuous product work creates stale tests, duplicate tests, and cases that still look right even though the business rule moved underneath them.

A version-aware loop that keeps the suite trustworthy

Compare the requirement revisions.
Start with the diff, not the old suite.
Look for changed rules, removed states, new actors, and new compliance constraints.

Identify impacted user flows and states.
Not every edit should trigger every test.
Target only the cases that the change touches.

Regenerate only the affected cases.
Keep stable cases stable.
That reduces churn and review fatigue.

Deduplicate against the existing suite.
A new case that repeats an old failure mode adds noise, not confidence.
If the path is already covered, link it instead of cloning it.

Report coverage by risk and version.
A suite linked to obsolete behavior is a liability.
Version-aware reporting keeps that visible.

Empirical work on agile testing has found that continuous requirement change increases test irrelevance and redundancy, and one prioritization model reported more than 90% improvement in reducing those problems plus an 80% earlier fault-detection rate in its evaluation (agile testing and prioritization study). Those figures aren't universal, but they point to the right operating model. The task is not to produce more tests. It is to keep the right ones alive.

The weekly ritual that prevents drift

A small ritual is enough if it is consistent.

  • Monday. Pick one PRD or user story set, then run the ambiguity classifier on every acceptance criterion.

  • Tuesday. Build the intermediate model, with states, decisions, actors, and outcomes.

  • Wednesday. Generate at least one positive, one negative, one boundary, and one recovery test per criterion, then assign requirement IDs.

  • Thursday. Compare the generated edge cases against the manual list from design review.

  • Friday. Publish the traceability matrix and run a coverage-gap review with QA.

That cadence does something subtle. It turns test generation into a requirements-quality diagnostic that happens every week, not just at release time. It also gives product and QA a way to ask whether the suite still matches the live business rule, which is the question that matters most when the product shifts under real users.

The grounded move for tomorrow morning is simple. Open the PRD that shipped last sprint, count the edge cases it never described, and use that count as the diagnostic for the next one. If you want a faster way to turn that habit into practice, visit Figr and use it to ground PRDs, flows, edge cases, and test cases in the actual product context your team already ships from.