What the work involves
You run adversarial conversations against AI models and agents, in both English and Malayalam, and write up what breaks. A typical block of work means picking a threat category from a provided taxonomy, constructing attack attempts against it — direct jailbreaks, indirect prompt injection, role-play framing, multi-turn escalation, bias elicitation, code-switching or transliteration tricks that exploit weaker non-English safety coverage — and then recording each attempt with enough precision that another person can reproduce it. Failed attacks matter too: coverage notes are part of the dataset. When a model does fail, you classify the vulnerability against the benchmark, judge severity, and flag whether it looks like a one-off or a systemic gap.
The project touches sensitive material: bias, misinformation, harassment and disinformation probing, and harmful-behavior elicitation. Mercor states that topics are communicated before exposure, that higher-sensitivity work is optional, and that wellness resources are provided. Take that at face value and decide honestly what you want to opt into.
What the screen looks for
- Real Malayalam depth, not conversational comfort. Screeners probe register, dialect, script versus Manglish transliteration, and whether you can judge if a model's Malayalam output is actually harmful or merely awkward.
- Structured adversarial method. The listing is explicit about frameworks and playbooks over "random hacks." Expect to be asked how you'd systematically cover a threat category rather than to recount one clever jailbreak.
- Reproducibility instinct. Can you describe an attack case the way a report would — setup, exact prompts, model response, severity, variants tried?
- Prior red teaming or adjacent work — AI adversarial testing, penetration testing, abuse and trust-and-safety analysis, disinformation research, or creative probing backgrounds.
Logistics
Fully remote and asynchronous, contract-based, paid hourly at an observed $20–22/hr — bands on Mercor vary by project and are not guaranteed. Work arrives in project batches, so volume fluctuates; people who can flex hours week to week get more of it. Expect assessment tasks before any paid work, and expect to be moved across customers and threat areas as projects rotate.