One of our design leaders handed an AI a Figma frame and asked for design critique. The results draw a precise map of where AI design review works today, and where it doesn't yet.
One of our design leaders ran the experiment several of us have been circling for months: hand Claude Code a Figma link and ask it to critique the design against Nielsen Norman heuristics. Not a demo, not a vendor pitch, an actual frame from an actual product, annotated by the model.
The results were useful in both directions. One frame took about 20 minutes and roughly $5 in tokens. The accessibility checks were genuinely good, because contrast ratios and touch targets are black-and-white rules, and rules are exactly what a model can verify at scale. The heuristic critique was slow and thin, because heuristics ask for judgment, and judgment needs context the model didn’t have. Who the user is, what the flow is trying to accomplish, what the team already debated and rejected 2 quarters ago.
Cheap experiments that fail precisely beat expensive plans that fail vaguely.
The finding generalizes, and it’s worth carving on the wall: AI critique is strongest where the standard is a rule and weakest where the standard is judgment. So the scoping is moving down rather than pushing harder. Point the tooling at design standards, the checkable stuff, and let Figma’s built-in design checks cover the smaller lint-level issues. Save the judgment for the crit room, where it belongs (this issue’s crit piece is a pretty good demo of what that room can do).
What I appreciated most is the economics of how we learned this. The experiment cost $5 and an afternoon, and it produced a precise answer about where the boundary sits today. Compare that to the alternative version of this story, a quarter-long tooling initiative built on the assumption that heuristic critique would scale, discovering the same boundary after 3 months and a lot of goodwill. Cheap experiments that fail precisely beat expensive plans that fail vaguely, and our leader’s willingness to share the failure openly is what made it valuable to all of us instead of just to him.
The timing matters too. A wave of new AI tooling seats is rolling out to our broader product team over the next couple of weeks. My hope is that we get a dozen more experiments like this one, small, honest, and shared, because that’s how a team actually learns a new medium: by finding its edges in public.
The teachable part
AI critique is strongest where the standard is a rule (contrast, touch targets) and weakest where the standard is judgment. A $5 experiment that finds that boundary beats a quarter-long plan that assumes it away.