What the work actually involves

You write prompts that a frontier model should plausibly fail, then judge whether it did. In practice that means moving between two modes: authoring questions of graded difficulty across fine arts, visual arts, architecture, design, art history, museums, and individual artists; and evaluating model output for factual accuracy, reasoning quality, completeness, and the kind of interpretive nuance art writing lives on. You will compare multiple candidate responses and rank them, explain in writing why a response is wrong with a citation an AI researcher can check, flag hallucinated attributions, invented provenance, misdated works, and confidently wrong iconographic readings, and rewrite ambiguous prompts that would produce unfair grading. Some of the work is adversarial by design — building benchmark items and edge cases where a model's plausible-sounding answer collapses under a specialist's scrutiny.

What the screen is looking for

Turing's screening process starts with an assessment you complete on your own, and the assessment rewards specificity over eloquence. Expect to be probed on where your art knowledge is genuinely deep versus where you'd defer, whether you can distinguish a contested scholarly position from a factual error, and whether you cite things a reviewer can verify — catalogue raisonné numbers, museum accession records, exhibition catalogues, peer-reviewed literature — rather than gesturing at general consensus. Writing that is precise and legible to a non-specialist reviewer matters more than writing that sounds curatorial. Prior LLM or prompt-engineering experience is preferred but not required; demonstrated evaluative judgment usually substitutes.

Credentials and logistics

  • Master's degree or higher in any field, with art-related disciplines preferred: art history, fine arts, visual arts, museum studies, cultural studies, design, humanities.
  • 3+ years of relevant professional work — art research, academia, museums, galleries, curation, criticism or journalism, education, publishing, or arts and cultural content.
  • Fully remote and contractor-classified: no medical or paid leave.
  • 40 hours per week for an eight-week contract, with a minimum of four hours daily overlapping US Pacific time.

Pay is not disclosed in this listing. Turing's expert-evaluation contracts are typically hourly and vary by domain and region; treat any figure you see quoted elsewhere as observed rather than promised, and confirm the rate in writing before you commit to the schedule.