What the work actually is

You write behavioral economics problems hard enough to break a frontier model, then judge whether its answers hold up. A task might ask the model to diagnose which bias explains a field-experiment result, redesign a default-enrollment scheme under a stated welfare objective, compute certainty equivalents under a prospect-theory value function with a given probability weighting, or spot the confound in an A/B test with differential attrition. Every item needs a defensible reference answer and a rationale that cites the mechanism, not just the label — "this is loss aversion" is not an answer key.

The grading half is equally weighted. Models are fluent about Kahneman and Thaler and often produce answers that sound like a textbook chapter while getting the direction of an effect backwards, invoking hyperbolic discounting where present bias is not identified, or treating a nudge result as robust when it comes from a single underpowered study. Your feedback has to name the specific error and point to the standard the answer violated.

What the screen looks for

  • Real research depth, not vocabulary. Expect follow-ups: which effects replicate, what the failure rates in the bias literature imply, when a framing manipulation is actually a change in information.
  • Grading discipline. Can you separate a wrong answer from a differently-argued right one, and hold a consistent line across dozens of items?
  • Item construction. Questions must be unambiguous, single-answer where claimed, and not solvable by pattern-matching on a famous study name.
  • Sourcing habits. Reference answers grounded in named papers and effect sizes rather than secondhand popularizations.

Logistics

Fully remote and asynchronous, contractor basis. Task batches are picked up on your own schedule; most contributors work in blocks of a few hours and commit to a rough weekly volume rather than fixed hours. Pay is hourly within an observed $60–100 band, tied to credentials and demonstrated task quality — not a guarantee. A short paid or unpaid calibration exercise before full onboarding is typical for this kind of engagement.