What the work involves
You run adversarial conversations against AI models and agents, working in both English and Kannada. In practice that means picking a harm category from a taxonomy, designing an attack path — roleplay framing, obfuscated instructions, code-switching between English and Kannada, instruction injection inside pasted content, slow escalation across multiple turns — and running it until the model either holds or fails. When it fails, you write it up: the prompt sequence, the exact output, the vulnerability class, a judgment on severity, and enough detail that someone else can reproduce the failure without guessing what you did. Kannada matters here because safety training is overwhelmingly English-first; the same request that gets refused in English often sails through in Kannada, or in Kannada written in Latin script, or when the harmful payload is Kannada wrapped in an English frame.
The project touches sensitive material — bias, misinformation, harassment, and content describing harmful behaviour. All of it is text. Topics are stated before you see anything, higher-sensitivity tracks are opt-in, and wellness resources are provided. This is worth thinking about honestly before you apply rather than after.
What the screen looks for
Mercor's process is AI-led interview plus profile review. Expect Kannada fluency to be tested directly, not asserted — likely including register shifts, regional vocabulary, and transliteration habits. Expect follow-ups on your red teaming background that ask for a specific attack you ran, not a category you are familiar with. The strongest candidates come with something concrete: jailbreak work on a named model family, pentesting or exploit experience, trust-and-safety abuse analysis, or disinformation research. Creative backgrounds — writing, psychology, acting, improv — count when you can show they translate into structured probing rather than scattershot poking.
Logistics
- Fully remote, asynchronous, project-based work assigned in batches.
- Observed pay band $20–22/hr; rates vary by project and are set by the platform, not guaranteed.
- Hours are flexible with no fixed shifts, but batches carry deadlines and throughput expectations.
- You will move between customers and taxonomies, so expect to re-read guidelines often.