Skip to content
Independent guides for QA & test automationRSSEditorial policy
QA Vibes

Free tool

Acceptance criteria grader

Paste a user story and see what a tester will have to ask you before they can write a single test. Ten checks, no sign-up, nothing leaves your browser.

Grading runs entirely in your browser. Nothing you paste is sent anywhere.

Scenario skeletons (4)
Feature: Acceptance criteria

  Scenario: Checkout should be fast
    Given TODO: the starting state for criterion 1
    When TODO: the action under test
    Then TODO: what can a tester see when “Checkout should be fast” holds?

  Scenario: Payment works correctly
    Given TODO: the starting state for criterion 2
    When TODO: the action under test
    Then TODO: what can a tester see when “Payment works correctly” holds?

  Scenario: The page is user-friendly on mobile
    Given TODO: the starting state for criterion 3
    When TODO: the action under test
    Then TODO: what can a tester see when “The page is user-friendly on mobile” holds?

  Scenario: Handle errors gracefully, etc.
    Given TODO: the starting state for criterion 4
    When TODO: the action under test
    Then TODO: what can a tester see when “Handle errors gracefully, etc.” holds?

Each criterion becomes one scenario. A TODO marks a gap the story never filled — a starting state or an action nobody wrote down. Those are the questions to bring to refinement, not for a tester to guess at later.

40

Score / 100

Will stall in refinement

  • Role, goal, and reason

    Partial · 5/10

    “user” is everyone, so it constrains nothing. Name the role whose day changes — the one whose edge cases you would think to ask about. Why this matters

  • Acceptance criteria listed

    Pass · 20/20

    4 criteria found.

  • Observable outcomes

    Missing · 0/15

    The criteria describe intent, not outcomes. Rewrite each one so it ends in something a tester can see or query. Why this matters

  • No vague wording

    Missing · 0/15

    5 vague terms (“fast”, “user-friendly”, “correctly”…). Each one is a question a developer will answer for you, in code, after the sprint. Why this matters

  • Failure and edge cases

    Missing · 0/10

    Only the happy path is described. Say what happens on invalid input, at the limit, and when the call fails. Why this matters

  • Measurable limits

    Missing · 0/10

    The story makes a performance or capacity claim (“fast”) with no number attached. Give it a threshold, or it can never pass or fail. Why this matters

  • One story, not several

    Pass · 5/5

    Scoped as a single story.

  • Starting state named

    Missing · 0/5

    Name the starting state: who is signed in, with what permission, and what data already exists. A tester who has to guess will test a different feature. Why this matters

  • Enough detail

    Pass · 5/5

    33 words: specific and still readable in refinement.

  • No secrets or real personal data

    Pass · 5/5

    No credentials or real personal data in the examples.

How scoring works

Having acceptance criteria at all is worth 20 points, because a story without them is a conversation someone will have to repeat. Observable outcomes and unambiguous wording are worth 15 each: those two decide whether a criterion can ever pass or fail. Failure and edge cases, measurable limits, and the story's role-goal-reason shape are worth 10 each; scope, starting state, length, and the check for leaked credentials are worth 5 each. A partial match earns half the points.

What “observable” means here. A criterion earns the point when it names something a tester could see, receive, or query without asking anyone: a message, a redirect, a disabled button, a saved record, a status code. “Payment works correctly” names none of those, which is why it scores nothing and why it survives refinement so often.

Why vague words cost so much. Words like fast, properly, user-friendly, and gracefully are questions in disguise. Left in a ticket, a developer answers them in code, and whether the answer was right is discovered in review or in production. A number or an exact message costs a minute now.

The Gherkin export. Each criterion becomes one scenario. Criteria already written as Given/When/Then keep their own steps; anything else becomes a skeleton with TODO where the story never said what the starting state or the action was. The TODOs are the point — they are the refinement questions, found before the sprint instead of during it.

The grader uses pattern matching, not AI, so treat it as a checklist rather than a verdict: it can be fooled by a story that says the right words about the wrong feature. The reasoning behind each check, and a story rewritten from 48 to 100, are in how to write acceptance criteria a tester can’t misread. For where the criteria stop and the team’s own quality bar starts, see acceptance criteria vs definition of done.