What the work involves

You are handed documents that plausibly belong in a data science workflow: an EDA notebook, a model card, an A/B test readout, a feature store schema doc, a post-mortem on a drifted model. Your job is to decide whether each one is realistic — not whether it is good, but whether a practitioner at a serious technology company would actually have written or received it. That means catching the tells: a train/test split described in a way no one would ever set up, a p-value quoted without a power calculation anywhere in sight, a notebook whose narrative order doesn't match how analysis actually unfolds, a model card with fairness sections that read like they were written by someone who has never had to fill one in.

Most tasks pair a judgment with a written rationale. The rationale is the product. "Unrealistic" is worth nothing to the lab; "unrealistic because the offline metric lift is reported to four decimal places on a 900-row holdout, which no one who has seen confidence intervals would do" is worth a great deal.

What the platform screens for

  • Recency and hands-on-ness. The listing asks for current familiarity with the documents your team writes and reviews. Expect the AI voice screen to push on what you shipped or reviewed in the last six months, not what you did in 2019.
  • Specificity under follow-up. Screeners re-ask. A general answer about "experimentation best practices" will be met with a request for the actual guardrail metric you used and why.
  • Calibration. Can you separate "this is unrealistic" from "this is realistic but bad"? Real practice is full of sloppy documents. Flagging every imperfection as fake is the most common failure mode.
  • Written output. Attention to detail is stated explicitly; expect a short written sample or a rationale-quality check.

Logistics

Fully remote and asynchronous. Commitment is flexible — contributors typically pick up batches rather than working fixed shifts, and volume moves with the lab's data needs rather than arriving at a steady rate. Rates of $150/hour have been observed on this listing; pay bands on evaluation work vary by task type and are never guaranteed. Expect a calibration period at the start where your labels are compared against a reference set before volume opens up.