What the work involves

You author self-contained behavioral science problems — a study design question, a measurement dilemma, an interpretation of a result set — and then write the reference answer that explains the methodological reasoning behind the correct conclusion. The reference answer matters more than the item itself: graders and models are both scored against it, so it has to name the assumption being tested, the inference that follows, and the inferences that do not. A second stream of work is grading: you read model output on existing items and rate it for construct validity, appropriateness of the statistical claim, and whether the model registers that a single significant finding is not a settled effect.

The failure mode you are hired to catch is fluent overreach. Models produce prose that sounds like a competent psychologist — correct vocabulary, plausible citations, confident effect language — while confusing a manipulation check with a mediation test, treating a convenience sample as generalizable, or reporting an effect size that the described design could not have estimated. Flagging those cases with a written rationale is the core deliverable.

What the platform screens for

AfterQuery's screen is AI-led and follow-up heavy. It will ask about specific studies you ran or analyzed and then push on the details — how the construct was operationalized, what the power analysis assumed, why you chose one estimator over another. Depth is measured by whether your answers survive three questions, not by the impressiveness of the first one. Expect scenarios where you must grade a plausible-but-wrong model response and articulate exactly which step broke.

Logistics

  • Fully remote and asynchronous; no standing meetings, no set timezone
  • Minimum commitment of 10 hours per week, with volume flexible week to week
  • Weekly payment via Stripe; the rate within the $60–120/hr band is set from credentials and demonstrated depth, and is observed rather than guaranteed
  • Work is routed toward your declared specialty — cognitive, social, clinical, or I/O