Introduction
This site publishes advice about test automation, so it should follow that advice. Every change runs through the kind of pipeline our CI/CD guide describes: lint, typecheck, a production build, then Playwright unit, end-to-end, and accessibility suites on desktop and mobile viewports.
During one build session, while we fixed rendering issues and added the bug report grader, that setup caught three real defects. None of them was exotic. All three had passed a visual review. This article walks through each one using the actual code, because the details matter more than the lessons.
The setup in one paragraph
Tests live in three Playwright projects: unit for pure functions (no browser), desktop and mobile for the production build served locally. The browser suites cover smoke checks on every route, article rendering, a crawler that follows every sitemap URL and internal link, SEO metadata, the grader, and an axe-core scan for WCAG 2.1 AA violations. GitHub Actions runs the same commands on every push:
- name: Build
run: npm run build
- name: Install Playwright browsers
run: npx playwright install --with-deps chromium
- name: Unit and end-to-end tests
run: npm testBug 1: every link color class was silently ignored
Symptom
The accessibility suite failed on the home page with one color-contrast violation. The axe report gave the exact colors: foreground #0f172a, background #005f5a, contrast ratio 2.36:1, against the required 4.5:1. The element was the "Open the grader" button, which the markup clearly styled as white text on dark teal:
<a class="... bg-teal-800 text-white ..." href="/tools/bug-report-grader">Open the grader</a>The button was not rendering white text. It showed dark slate text on dark teal, which is close to unreadable.
Why review missed it
Other links styled with text-white, such as the hero buttons, looked correct. Their parent section happened to set white text too, so the links inherited the right color by accident. The bug only showed on a link whose parent used a different color.
Root cause
The global stylesheet contained a harmless-looking rule outside any cascade layer:
a {
color: inherit;
}Tailwind CSS v4 puts its utilities inside @layer utilities. In the CSS cascade, styles outside any layer beat styles inside layers, regardless of specificity. So this one-line element selector overrode every text-* utility on every link on the site.
Fix and regression test
Base element styles moved into @layer base, and the redundant a rule was removed because Tailwind's preflight already sets it inside a layer:
@layer base {
body {
background-color: var(--background);
color: var(--foreground);
font-family: var(--font-sans);
}
}A regression test now checks the computed color, not the class name:
test("text color utilities apply to links (regression: unlayered base CSS)", async ({ page }) => {
await page.goto("/");
const cta = page.getByRole("link", { name: "Open the grader" });
await expect(cta).toHaveCSS("color", /rgb\(255, 255, 255\)|oklch\(1 0 0\)/);
});Lesson: a class in the markup does not guarantee the style applies. Automated contrast checks measure the colors that actually render, which makes them good at catching cascade bugs that screenshots and code review both miss.
Bug 2: the grader approved its own blank template
Symptom
The bug report grader scores a pasted report against a set of weighted checks. One unit test was written as an adversarial case before any bug was known: feed the grader its own empty template and make sure the result is not "Ready to file". That test failed.
test("unfilled template does not score as ready", () => {
expect(gradeBugReport(TEMPLATE).grade).not.toBe("Ready to file");
});Root cause
Several checks looked for keywords. The template is full of the right keywords with nothing after them: headings like ## Environment and ## Expected, and labels like Severity: and Build / version:. The environment check saw "Build" and "Browser", the impact check saw "Severity", and the expected-vs-actual check saw both headings. A report that said nothing scored like a good one.
Fix
Keyword checks now run on a copy of the text with the empty scaffolding removed, and section checks require real content under the heading:
function stripPlaceholders(text: string): string {
return lines(text)
.map((l) => l.replace(/<[^>\n]{1,40}>/g, ""))
.filter((l) => !/^\s*(?:[-*+]\s*|#{1,6}\s*)?[^:\n]{1,60}:\s*$/.test(l))
.filter((l) => !/^\s*(?:\d+[.)]|[-*+])\s*$/.test(l))
.join("\n");
}That removes empty Label: lines, empty list items, and <angle-bracket> placeholders. A companion function, hasFilledSection, only accepts a heading if at least one non-placeholder line appears before the next section.
Lesson: heuristic scoring needs negative fixtures. The most useful one was the tool's own template: it has the right structure and zero information, which is exactly the kind of input keyword matching gets wrong.
Bug 3: a sentinel value that collided with real data
Symptom
The grader gained a "Copy as Jira markup" button. The Markdown-to-Jira converter protects inline code by swapping it for placeholder markers, formatting the rest, then restoring the code. The markers are meant to be control characters (U+0001 to U+0003), which never appear in pasted text.
A file-editing step wrote them as the literal digit strings "0001", "0002", and "0003" instead. All 22 unit tests still passed.
How it was caught
Not by a test: by reading the diff and dumping the bytes of the changed lines. The tests passed only because no fixture contained those digits. With the broken markers, the final step that restores bold text would have replaced every 0003 in a report with an asterisk, so a build number like 2026.0003 would have been copied into Jira as 2026.*.
Fix and regression test
The markers are now written as explicit escape sequences, so they stay visible in source and cannot be mangled invisibly:
const CODE_OPEN = "\u0001";
const CODE_CLOSE = "\u0002";
const BOLD = "\u0003";And a test now feeds the converter the exact values that would collide:
test("digit sequences are never mistaken for internal placeholders", () => {
const text = "Build 0001-0002-0003 with `code 0001` and **bold 0003**";
expect(markdownToJira(text)).toBe("Build 0001-0002-0003 with {{code 0001}} and *bold 0003*");
});Lesson: a green suite proves only what its fixtures exercise. When code uses sentinel values, write a test whose input contains those sentinels.
What the three bugs have in common
- Each passed a visual check. The page looked fine, the grader produced plausible scores, and the converter handled every example we tried.
- Each was caught by a check that measured the output directly: rendered colors, a score for a known-bad input, and bytes on disk.
- Each fix shipped with a test that fails if the bug returns, named after the bug so the next person knows why it exists.
Conclusion
None of these bugs needed a sophisticated tool to find. They needed an accessibility scan on real pages, one deliberately hostile unit test, and a habit of looking at actual bytes instead of trusting what the editor shows. If your team already runs Playwright in CI, adding an axe scan and a few adversarial fixtures costs an afternoon, and in our case it found three defects in a single session.