What the work involves

You build the material AI labs use to measure design judgment. That means writing a realistic product brief, taking it through to a polished UI yourself, then documenting the reasoning: why this hierarchy, why this spacing rhythm, why this component instead of a custom one. A second stream is grading — a model produces an interface or a critique, and you score it against usability, information hierarchy, and visual craft, explaining what a working designer would flag and why. A third is Figma workflow capture: showing how auto layout, variants, or a handoff actually get used in practice, not how a tutorial says they should.

The hard part is almost never producing the design. It is articulating the tacit judgment behind it in language precise enough that a grader — human or model — can apply the same standard tomorrow. "Feels cluttered" is useless. "Three competing primary actions in the same visual weight; the destructive one should recede" is usable.

What the platform screens for

  • Shipped work, not concepts. A portfolio of real product or interface work you can talk through — constraints, tradeoffs, what changed after it went live.
  • Depth under follow-up. Expect the screener to push past your first answer on a specific design decision. Vocabulary alone doesn't survive two probes.
  • Written English that carries reasoning. Most of your output is prose, not pixels.
  • Tool fluency. Figma or a comparable product design tool, at the level of someone who maintains files other people work in.

Design systems, component library work, UX research breadth, and front-end or handoff familiarity (HTML/CSS) are preferred, not required.

Logistics

Fully remote, fully async — no standing meetings, no fixed timezone. Minimum commitment is 10 hours per week, and volume scales up or down between weeks. Pay is weekly via Stripe at a rate set from your experience within the observed $60–120/hr band; rates are assigned, not guaranteed, and vary by task type and seniority.