What the work involves
CUA data is agent-trajectory data: screenshots, clicks, keystrokes, tool calls, and the reasoning that connects them. As a Quality Analyst you sit downstream of the annotators producing it and upstream of the labs consuming it. Day to day that means sampling annotated tasks and scoring them against a rubric, writing feedback that a contributor can actually act on, and escalating the cases where the rubric itself is the problem rather than the annotator.
You are also expected to build, not just enforce. Expect to draft and revise annotation guidelines, run onboarding sessions for new annotators, participate in calibration exercises where reviewers score the same items and reconcile disagreements, and support pilot runs where the workflow is still being shaped. Turing describes this as a "functional pod" — you work alongside QA leads and project managers rather than in isolation, and edge-case resolution is collaborative.
What the platform screens for
Expect probing on three things. First, concrete prior exposure to annotation or evaluation operations — what rubric you worked against, what the quality bar was, how disagreement was measured. Second, judgment: given an ambiguous trajectory, can you say what you'd mark and why, and can you tell a guideline gap apart from an annotator error. Third, logistics, which are unusually rigid here: 40 hours a week for two weeks with four hours of daily Pacific overlap, plus a machine with at least 16GB RAM and a stable connection. Vague availability answers are a common cause of rejection on short-duration pods.
Logistics
- Remote contractor assignment, task-based, no paid or medical leave.
- Stated duration: two weeks, full-time at 40 hours/week.
- Four hours per day must overlap Pacific Time.
- Pay is undisclosed on this listing; ask for the rate and the payment cadence before accepting.
- Shortlisted candidates receive a Job Interest Form before any offer stage.