Skip to content
Independent guides for QA & test automationRSSEditorial policy
QA Vibes

How to Write Acceptance Criteria a Tester Can't Misread

Acceptance criteria decide what gets built and what gets tested. This guide covers the eight properties that make a criterion checkable, with a real story rewritten from a score of 48 to 100 in our acceptance criteria grader, and the questions each rewrite answered before the sprint instead of during it.

QA Vibes EditorialPublished Updated 11 minRevision history ↓

Key takeaways

  • A criterion earns its place only if someone who has never seen the feature can tell whether it holds.
  • Words like fast, properly, and user-friendly are questions in disguise; a developer will answer them in code.
  • Most escaped defects live on the paths a story never described: empty results, expired data, the service being down.
  • Rewriting one vague story took it from 48 to 100 in the grader and answered five questions before a line was written.
Contents (16 sections)

Introduction

Acceptance criteria are the conditions a story has to meet before anyone agrees it is finished. They are written before the work starts, and they decide two things at once: what a developer builds, and what a tester checks.

That double duty is why vague criteria are expensive. "Payment works correctly" does not describe anything a developer can build to, or a tester can check against. It survives refinement because everyone reads it and pictures something — the trouble is that they each picture something different, and the difference surfaces in review, or in production.

This guide covers the eight properties that make a criterion checkable. At the end, one real story is rewritten from a score of 48 to 100 in our acceptance criteria grader, and every point it gained corresponds to a question answered before the sprint rather than during it.

What acceptance criteria are for

A criterion has done its job when a tester who has never seen the feature can decide, without asking anyone, whether it holds.

That is the whole test. It rules out criteria that describe intent ("the filter should be useful"), criteria that describe implementation ("the query uses an index"), and criteria that describe a feeling ("the page feels responsive"). It keeps criteria that describe an outcome someone can observe: a message, a redirect, a disabled button, a saved record, a status code.

Most stories need at least three criteria: the main path, a failure, and a boundary. A story with one criterion is usually a story whose author only pictured it working.

Who writes them

The product owner starts. They know what the customer needs and why the work is worth doing, and the Scrum Guide makes them accountable for Product Backlog items being clear and understood — including delegating the writing, while keeping the accountability.

But criteria written alone tend to describe only the path their author imagined. The common practice is the Three Amigos: the product owner, a developer, and a tester write them together. Each brings a different blind spot. The product owner knows the rule; the developer knows which states the system can actually be in; the tester knows which of those states nobody has thought about.

Fifteen minutes of that conversation is cheaper than a sprint spent building the wrong reading of one sentence.

Name the role, the goal, and the reason

The familiar template earns its place when all three parts are specific:

As a support agent handling refund calls, I want to filter the order list by status and date so that I can find a caller's order while they are still on the line.

"As a user" is the version to avoid. It constrains nothing, because every person who touches the system is a user. Name the role whose day changes, and the edge cases start suggesting themselves: a support agent is on a call, under time pressure, looking at somebody else's data, and probably has permissions a customer does not.

The reason matters as much as the goal. When a trade-off comes up mid-sprint — and one always does — the reason is what the team uses to decide. "So they can find it while the caller is on the line" tells you that speed matters more than a sophisticated filter builder. Without the reason, that decision gets made by whoever is closest to the keyboard.

Make every criterion observable

Rewrite each criterion so it ends in something a tester can see, receive, or query.

Not observable Observable
Filters work correctly Only refunded orders placed in the selected range are listed, and the result count is displayed above the table
Empty results are handled The table is replaced by the message "No orders match these filters." and a Clear filters button is displayed
The user is notified A confirmation email is sent to the account address within 1 minute
Invalid input is rejected The field shows "Enter a date in the past" and the form is not submitted

The right-hand column is longer, and that is the point: the length is the information that was missing. Each one also happens to be a test case that can be written without another meeting.

A useful check while drafting: if a criterion could be pasted into a bug report as an expected result, it is observable. If it would have to be translated first, it is not finished.

The words that commit to nothing

Some words look like requirements and commit to nothing at all. They pass refinement precisely because nobody disagrees with them.

  • fast, quick, responsive, snappy — how many seconds?
  • properly, correctly, appropriately, as expected, gracefully — what is the right result?
  • user-friendly, intuitive, seamless, clean, modern — what does the screen show?
  • robust, scalable, reliable, secure — measured against what threshold?
  • several, some, many, various — how many?
  • etc., and so on — which cases, exactly?
  • as needed, if necessary, where applicable — when, exactly?

Each is a question in disguise. Left in the ticket, a developer answers it in code, and whether the answer matched what the product owner had in mind is discovered in review at best, in production at worst. Answering it while writing the story costs a minute.

The one to watch hardest is gracefully, as in "handle errors gracefully". It is the phrase that most reliably means nobody has decided what happens when the call fails.

Write down the failures too

The happy path is the part everyone pictures, so it is the part that gets written down. Most escaped defects live everywhere else.

For each story, ask these five questions and write down the answers that apply:

  1. Empty. What does the screen show when there is nothing to show?
  2. Invalid. What is the exact message, and what is left unchanged?
  3. Not allowed. What happens when the user does not have the permission?
  4. Too much. What happens at the limit — the maximum length, the largest file, the ten-thousandth row?
  5. Broken. What happens when the call fails or the service is down? Is the previous state kept?

Answering question 5 in the story rather than at build time is often the difference between a page that keeps the user's work and one that throws it away.

Put a number on every limit

If a story makes a claim about speed, size, or capacity, the claim needs a number, or it can never pass or fail.

  • "The list updates quickly" → "The filtered list is displayed within 2 seconds"
  • "It should handle large accounts" → "With 10,000 orders in the account, any filter returns within 2 seconds"
  • "The upload should be reasonable" → "Files up to 20 MB are accepted; larger files show 'Maximum file size is 20 MB'"

The number does not have to be right the first time. It has to be written down, because a number can be argued with, measured, and revised. "Quickly" cannot: it is only ever discovered to have been wrong.

A story with no limit at all is fine — plenty of stories have none. The mistake is having one and leaving it unstated.

One story, not several

Two signs that a ticket contains more than one story: a second "As a … I want …", and a criteria list that keeps going past a dozen.

Split it. Each part should be shippable on its own, which is the test that matters — not size, but whether half of it could go to production and be worth something. "Filter the order list" and "export the filtered list" are two stories; the first is useful without the second.

Anything genuinely out of scope is worth stating in the story, in its own short section. It is faster to write "exporting the filtered list is a separate story" than to have the same conversation in three standups.

Name the starting state

A criterion that does not say what was already true is a criterion two people will test differently.

Name the account and its permissions, the data that already exists, and anything that has to have happened first. "Given I am signed in as a support agent and the account has at least 20 orders" tells the tester which account to use and which data to seed. Without it, someone will test with three orders, the pagination criterion will pass, and the bug will be found by a customer.

This is also what the Given in Given/When/Then is for, which is the main argument for the format: it has a slot for the state, so the state is harder to forget.

Keep credentials out of the ticket

Backlog tickets are readable by the whole company, often by contractors and vendors, and they get exported into reports, screenshots, and support threads. They are not a place for a password, an API key, a token, or a real customer's email address or card number.

Use a placeholder, or an address on a reserved domain such as @example.com. If a real credential has already gone into a ticket, rotating it is the fix; deleting the comment is not, because the notification email has already been sent.

Bullets or Given/When/Then?

Both work. The difference is which mistakes each one makes harder.

Bullets are faster to write and easier to skim, and suit stories that are mostly a list of rules ("the export includes columns A, B, and C"). Their weakness is that the starting state is easy to leave out.

Given/When/Then forces a slot for the state, the action, and the result, which is exactly the shape a test needs. It is heavier to read in bulk, and it tempts teams into writing implementation steps ("When I click the button with id #filter") rather than behavior.

Pick one per story rather than mixing them mid-list. If the team uses Cucumber or SpecFlow, Given/When/Then criteria become executable specifications directly — see Cucumber and Gherkin by example for what that costs and what it buys.

Worked example: from 48 to 100

Here is a story as it arrived, scored in the acceptance criteria grader:

As a user, I want to filter the order list so that I can find orders faster.
 
Acceptance criteria:
- Filters work correctly
- The list updates quickly
- Handle empty results properly

Score: 48 — will stall in refinement. The grader flags no observable outcome in any of the three criteria, "faster" and "correctly" as vague, a speed claim with no number, no starting state, and a role that names everyone.

None of that is a style complaint. Each flag is a question someone will have to ask before they can write a single test: Which orders? Filtered by what? How quickly? What does "empty results handled" look like on screen? Which account, with how much data?

Answering those five questions in the story produces this:

As a support agent handling refund calls, I want to filter the order list by status and date so that I can find a caller's order while they are still on the line.
 
## Acceptance criteria
 
- Given I am signed in as a support agent and the account has at least 20 orders
  When I select status "Refunded" and the date range 1-31 August
  Then only refunded orders placed in that range are listed, and the result count is displayed above the table
- Given a filter combination that matches no orders
  When the filter is applied
  Then the table is replaced by the message "No orders match these filters." and a Clear filters button is displayed
- Given a filter is applied
  When I reload the page
  Then the same filter is still applied and the URL contains it
- Given the order service is unavailable
  When I apply a filter
  Then the list I already had is kept and the message "Filters are unavailable right now. Try again." is displayed
- Given an account with 10000 orders
  When I apply any filter
  Then the filtered list is displayed within 2 seconds
 
## Out of scope
 
Exporting the filtered list; that is a separate story.

Score: 100 — ready to build.

The rewrite is longer, but nothing in it is decoration. Two criteria did not exist before: the service being unavailable, and the behaviour at 10,000 orders. Both were discovered by asking the questions, not by testing — which is the cheapest place to discover them. The reload criterion came from the reason: an agent on a call who loses their filter on refresh is back where they started.

A checklist before refinement

  • The role is a specific person, not "a user", and the reason explains why the work is worth doing.
  • There are at least three criteria: the main path, a failure, and a boundary.
  • Every criterion ends in something a tester can see, receive, or query.
  • No vague words: fast, properly, user-friendly, gracefully, several, etc.
  • Every speed, size, or capacity claim has a number.
  • The empty case, the invalid case, the permission case, the limit, and the failure are answered or explicitly out of scope.
  • The starting state names the account, its permissions, and the data that must already exist.
  • It is one story; anything else is split out or listed as out of scope.
  • No passwords, tokens, keys, or real customer data.

Conclusion

Acceptance criteria are not documentation written after a decision; they are where the decision gets made. Every vague word left in a story is a question that will be answered anyway — by a developer at build time, by a tester at review time, or by a customer in production. The only variable is how expensive the answer is when it arrives.

The properties that make criteria checkable are mechanical enough to grade, which is why we built the acceptance criteria grader to do exactly that: paste a story, see what a tester would have to ask before writing a single test. What the criteria then become is covered in how to write effective test cases.

Sources and further reading

Revision history

Updated source links that had moved: the pages still exist, at new addresses.

Spotted a mistake? Report it — corrections land here.

Written and reviewed by

QA Vibes Editorial

Articles are written and reviewed by practicing QA and automation engineers. Every article lists its sources and shows when it was last updated.

Try it on a real story

Acceptance criteria grader

Paste a user story, see which criteria a tester could misread, and export them as Gherkin.