Introduction
"Manual or automated?" is usually argued in the abstract. We tried both on the same product and compared what each one noticed.
The product was Sauce Demo, a public practice shop from Sauce Labs. It has several accounts, and all but one are deliberately broken in different ways, which makes it a fair test: there are real bugs to find, and nobody has to invent them.
- Automated: a Playwright suite with a checkout journey, required-field checks, and a few targeted tests.
- Looking: opening the same pages as each account and comparing them with the working account, the way a manual tester would.
What automation caught, and what it missed
The checkout test logs in, adds a product, fills in the form, and checks the confirmation heading. We ran it for different accounts:
| Account | Checkout test result | What was actually wrong |
|---|---|---|
standard_user |
Passed | Nothing |
visual_user |
Passed | Prices were wrong, and different on every page load |
problem_user |
Failed at checkout | Typing a last name overwrites the first name |
error_user |
Failed at checkout | Checkout can't be finished; sorting shows an error alert |
The visual_user row is the interesting one. The checkout test passed while the product list showed prices like $82.16 and $72.42 for items that cost $29.99 and $9.99. The test never looked at prices, so it had no way to know.
A separate automated test that compares displayed prices with the catalog did fail, and so did a test checking that product images are distinct on problem_user, where every product shows the same placeholder picture. But we only wrote those two tests after we had already seen both problems while exploring the shop. Nobody writes an assertion for a bug they haven't imagined.
What looking at the pages reveals
None of these needs an assertion. Open the pages as each account, compare them with standard_user, and they're hard to miss:
- Every product image is the same placeholder on
problem_user. - Prices don't match the working account on
visual_user, and they change when you reload the page. - An alert says "Sorting is broken!" on
error_user, and the sort order doesn't change. - Login is slow on
performance_glitch_user: about five seconds before the product list appears. - The last name field behaves strangely on
problem_user: typing in it changes the first name.
A person needs no assertion to notice that something looks wrong or feels slow. That's the core strength of manual testing, and exactly what a script lacks.
What a person would miss
The comparison cuts both ways. The same person would not reliably:
- Repeat the checkout for every account on every release. The automated suite did it in seconds; a person doing it by hand gets bored, skips steps, and misses the day it breaks.
- Notice that a total changed from
$32.39to$32.40. The automated assertion checks the exact value every time. - Test three required fields with the same care on the fiftieth run. Our data-driven test checks each one, with the exact error message, on every run.
- Catch a slow creep in response times from one release to the next without measuring them.
The pattern
| Manual testing | Automated testing | |
|---|---|---|
| Finds | Things that look or feel wrong, confusing flows, unexpected behavior | Regressions in behavior someone already specified |
| Needs | A curious person and time | A clear expected result written as an assertion |
| Cost per run | The same every time | Near zero after it's written |
| Cost of change | Low: update notes | Medium: update code |
| Misses | Repetition, precision, scale | Anything nobody thought to assert |
Automation checks what you already know should be true. Manual testing, especially exploratory testing, discovers what you didn't know to check. In our run, every automated test that found a visual or data bug was written because a person had found it first.
Deciding what to automate next
For each check you're considering, ask:
- Will this run many times? Checks repeated every release or every pull request are the best candidates.
- Is the expected result exact? "Total is $32.39" can be automated. "The page looks right" mostly can't, although visual comparison tools cover part of it; see our visual regression testing guide.
- Is it stable? A feature redesigned every week makes automated tests expensive. Test it by hand until it settles.
- What does a miss cost? Checkout, login, and payments justify automation even when change is frequent.
- Can it run below the UI? A pricing rule checked through the API is faster and less fragile than the same rule checked through the browser. See the test pyramid.
When the answers are "rarely", "judgment", "changing", and "low impact", keep it manual.
How the two work together
A workflow that uses both, based on what happened here:
- Explore new or changed features by hand first. Use a short charter: "Explore checkout as
problem_userto find input handling problems." - Turn every bug found into a precise, automated check where the expected result is exact, so it can't come back unnoticed.
- Keep a small automated suite on the journeys that earn money, and run it on every change.
- Spend the time automation saves on more exploration, not on writing more scripts for their own sake.
Conclusion
In our comparison, automation passed a checkout with wildly wrong prices, and a person spotted the problem in seconds. The same person could never have re-checked every account, field, and total on every release the way the suite did. The question isn't which one to choose. It's which checks to hand to code, so people have time for the exploration no script can do.