What the work involves
This is a compressed engagement, not an ongoing contract. Over two consecutive days you will work through a queue of safety-relevant tasks — typically probing a model for unsafe or policy-violating outputs, writing adversarial prompts, judging whether a given response crosses a stated line, and documenting what you found in clear, reproducible prose. Sprints of this shape usually pair each judgment with a short written rationale, so the value of your work sits as much in the explanation as in the label. Expect rubrics that arrive with the task rather than in advance, and expect them to be refined mid-sprint as the client sees early submissions.
What the platform screens for
- Genuine adversarial instinct. Screeners want to hear how you actually construct an attack — role-play framing, incremental escalation, obfuscation, multi-turn setup — not a recitation of jailbreak headlines.
- Calibrated severity judgment. Distinguishing a harmful output from a merely distasteful one, and knowing which policy line a response crosses and why.
- Written precision. Rationales need to be readable by someone who never saw the conversation. Vague, hedge-laden writing is the most common reason work is rejected.
- Hard logistics. U.S. location and blocked availability for both days are checked directly; partial availability is generally a decline rather than a negotiation.
Logistics
Fully remote and asynchronous within each day, but not open-ended: roughly five hours per day for two consecutive days during the week of September 1, 2026, for about ten hours total. Pay is observed at $45 per task rather than hourly, so throughput and first-pass quality both matter — reworked tasks cost you time you are not compensated for. You will need reliable internet and the ability to respond to clarifying messages during the sprint window, since rubric updates often land mid-stream.