What the work involves

You sit on the adversarial side of the table. Day to day you construct prompts intended to make a model fail — jailbreak attempts, indirect prompt injection, misuse framings, bias elicitation, and multi-turn manipulation where the payload is built up across several conversational turns rather than stated outright. A substantial portion of this project is cross-lingual: attacks that a model's English safety training catches may pass unfiltered in Telugu, and code-mixed Telugu-English (including romanised Telugu) is a common vector. Native command of both languages is what makes you useful here, not a bonus.

The second half of the job is documentation. An unrepeatable jailbreak has little value to a customer. You annotate what failed, classify the vulnerability against the project's taxonomy, note the exact prompt sequence and model conditions, and write it up so an engineer who does not speak Telugu can understand the risk and reproduce the case. Expect to follow playbooks and benchmarks rather than freelance — consistency across annotators is what makes the dataset usable.

What the screen looks for

Mercor's screening is AI-led and follow-up heavy. It probes for concrete prior adversarial work — AI red teaming, penetration testing, disinformation or abuse analysis, trust and safety investigation — and it will push on specifics: what you tried, what the model did, why you thought that vector would work. Generic enthusiasm for "breaking things" does not survive two follow-ups. It also tests Telugu depth in ways that matter for this project: register, dialect variation, script versus romanisation, and whether you can judge that a Telugu output is harmful in context rather than merely literally translated.

Logistics and content exposure

  • Fully remote, asynchronous, task-based volume rather than fixed shifts
  • Text-only work; no audio, video, or imagery
  • Project topics touch bias, misinformation, and harmful behaviours; topics are disclosed before exposure
  • Higher-sensitivity tracks are optional, with stated guidelines and wellness resources
  • Pay observed at $20–22/hr for this listing; bands vary by project and are not guaranteed