Screenshot a page, ask “what do you think of this design”, and you will get polite praise and generalities: generous whitespace, clear hierarchy. None of it is wrong, but it could be said about any screen at all, which makes it useless.
The problem is the question, not the tool. Throw a mockup at a human designer with no context and you get the same answer. A useful critique only arrives once you have said what to judge it against.
Fill in four boxes
Rather than writing a long paragraph, think of it as filling four boxes. Do that and the quality stays consistent from one review to the next.
Context reads like “a service page seen by a first-time visitor arriving from search”. Audience reads like “a small business owner running their own WordPress site, deciding here whether to enquire”. With those two present, the critique shifts from “is this attractive” to “does this serve the purpose”.
The criteria box matters most. Without a fixed list, each review looks at something different and none of them compare with the last. Fix the list and the same yardstick gets applied repeatedly — which is, in effect, your team’s design rules.
The format box makes the answer actionable. Ask for “a severity rating and a concrete change for each finding” and instead of “consider adjusting the spacing” you get “the gap between cards is the same as the gap between sections, so the grouping does not read — keep the card gap and increase the section gap”.
What it catches, and what it cannot
Knowing the split saves a great deal of time. An AI critique is good at rule violations and omissions. Where judgement is required, it produces plausible sentences and little else.
Do not let it eyeball contrast in particular. Given a screenshot it will happily estimate the colours and pronounce them “probably sufficient”, but accessibility is a calculated ratio, not an impression. Give it the actual colour values and ask for the arithmetic, or better, measure with a tool.
Move repeated findings into rules
When the same finding appears three times, it is not a review item — it is a rule you have never written down. It means the spacing scale was never fixed, or the button hierarchy was never documented. Consume critiques only as individual fixes and the next screen earns you the same finding again.
One caution: these models lean towards agreement. Ask “isn’t this version better?” and the answer is usually yes. When comparing two options, do not reveal which one you prefer — ask separately for the weaknesses of each. The quality of the answer changes noticeably.
Other ways of wiring AI into a working process live in the AI archive, and turning recurring accessibility findings into rules is covered in the accessibility series. The order in which we review and verify our own work is published on the process page.