What the work involves

You read conversations in which a person brings a real interpersonal or everyday decision — a family conflict, a workplace dilemma, a health-adjacent choice, a financial or educational tradeoff — to an AI model, and the model responds with advice. Your job is not to grade the advice on style. It is to ask what the model should have known before answering: which facts were missing, which follow-up question would have surfaced them, and whether knowing them would actually have changed the recommendation. That last test is the crux of the role. Turing's guidelines frame it as separating "load-bearing" context from information that is merely interesting, and much of your written output will be a defensible argument for which side a given gap falls on.

Day to day this looks like structured qualitative review: read the transcript, tag missed elicitation opportunities, draft the follow-up questions you would have asked in the model's place, rate whether the recommendation was safely groundable on what the model actually knew, and write rationales that a second reviewer could apply consistently. Expect calibration sessions, rubric updates mid-project, and disagreement resolution with research staff. The work rewards people who already interview for a living — intake assessors, therapists, mediators, user researchers, academic advisors — because the underlying skill is knowing which unasked question would have changed the plan.

What the screen looks for

  • A verifiable credential in one of the four named tracks — licensed social work (MSW plus LCSW/LICSW/LMSW), counseling or psychology (licensed master's, or PhD/PsyD in clinical, counseling, social, or I-O psychology), advanced qualitative/HCI/communication research, or demonstrated advisory expertise in mediation, coaching, student services, or organizational consulting.
  • Elicitation experience you can narrate concretely — screeners push for a specific case where a question you asked changed your assessment, and for the reasoning behind asking it.
  • Rubric discipline — whether you can suppress your own clinical preference when the guideline says otherwise, and flag the guideline rather than quietly deviating.
  • Honest availability — 40 hours a week for four weeks with four hours of PST overlap is a real constraint, not a target.

Logistics

Remote and contractor-based, with no medical or paid leave. The stated commitment is 40 hours weekly across a four-week contract, with at least four hours of daily overlap with Pacific time for calibration and reviewer sync. Payment is pay-per-task; Turing has not disclosed a rate for this listing, and throughput on task-based evaluation work varies with rubric complexity, so treat any rate you are quoted as project-specific rather than a platform standard. Contracts of this shape are sometimes extended, but the posted duration is four weeks and you should plan around that.