Change one thing in CSS and something else changes with it. That is both its strength and its hazard. Adjust the spacing on a shared component by one step and two of the twelve screens using it will shift — this happens constantly.
The difficulty is that the person who made the change looks only at the screen they were changing. Nobody opens the other eleven, and a few days later a visitor finds it first.
What visual regression actually catches
The mechanism is simple. Store an image of each screen at a known-good moment, capture the same screens after a change, and compare them pixel by pixel. Anything past a threshold gets reported.
One thing is easy to misunderstand. This check does not decide whether the design is good; it decides only whether it changed. So it catches exactly one class of problem — unintended change. Intended changes have to be approved by a person and promoted to the new baseline, and that approval step is the real running cost of adopting the tool.
It earns that cost because unintended change is the slowest to discover and the hardest to explain. Almost every “how long has it looked like this?” after a release comes from this category.
Choosing what to guard is half the job
Capturing every page makes the run slow and, worse, produces a diff report long enough that nobody reads it. The moment the list gets long, the tool is finished.
The criterion is what it costs when it breaks. Revenue paths and shared components come first; frequently changing content pages come last.
A particularly useful trick is to build a page that gathers the components and capture that. Lay out buttons, cards and form elements in every state on one page and a single image guards the whole design system against regression.
Suppressing false alarms is the other half
Adoption nearly always fails the same way: everything differs every time, so nobody looks any more. The causes are predictable enough to head off in advance.
Fonts are the usual culprit. Capture before the web font arrives and everything is drawn in the fallback at different widths, so the entire screen reports as changed. Simply waiting for font loading to settle removes a large share of the noise.
Animation is the same story: capture while an entrance animation is running and you catch a different frame each time. Turning motion off for the duration of the capture yields the static final state — and the fact that this works at all is itself evidence that the screen was built complete without motion.
Tune the threshold too. Near zero and you catch sub-pixel font rendering differences; too high and you miss real breakage. Starting slightly loose and tightening as you eliminate false alarms is the practical route.
What this check cannot replace
Be clear about the limit. Screenshot comparison sees only what is visible. What a screen reader announces, whether keyboard order makes sense, whether contrast clears the threshold — none of that appears in an image.
So it does not replace accessibility testing or human review; it holds the line those two already established. The standard comes first, and this check keeps it.
Settling screen states and rules during review is covered in the workflow archive, and everything that does not show up in an image lives in the accessibility series. The order in which we verify our own work is published on the process page.