What the work actually involves
You run adversarial conversations against AI models and agents in both English and Tamil, then write up what broke. A typical session means picking a risk category from a supplied taxonomy — misinformation, harassment, bias exploitation, unsafe instruction-following — and constructing attacks that get the model to produce output it was trained to refuse. That includes direct jailbreaks, indirect prompt injection through content the model reads, role-play and persona framing, and slow multi-turn manipulation where the payload only appears on turn six. The Tamil side matters because safety training is usually thinnest outside English: a refusal that holds firmly in English often fails when the same request arrives in Tamil, in Tanglish transliteration, in code-switched Tamil-English, or wrapped in regional idiom and honorifics. Finding and documenting those gaps is much of the value here.
The second half of the job is discipline. An attack that works once and can't be reproduced is worth little. You label the failure, classify the vulnerability against the project's schema, record the exact prompt sequence and conditions, and estimate severity and generality — does this work only on one phrasing, or does the whole family of attacks land? Reports go to customers who need to act on them, so writing clearly for both engineers and non-technical stakeholders is part of the deliverable, not an afterthought.
What the screen looks for
- Genuine native-level Tamil and English, including registers — formal written Tamil, colloquial speech, transliterated Tanglish, and code-switching. Expect probing beyond "are you fluent."
- Prior adversarial work of some kind: AI red teaming, penetration testing, exploit development, trust-and-safety abuse analysis, or disinformation and harassment research. Creative-probing backgrounds — psychology, acting, writing — count when paired with structured method.
- Method over improvisation. Screeners want to hear that you work from taxonomies, benchmarks, or attack playbooks and can say why one technique was chosen over another.
- Reproducibility instincts. Can you describe a finding so someone else can trigger it on demand?
- Calibration on harm. Knowing what should be refused, what is a false refusal, and where the line sits in Tamil-language cultural, caste, religious and political context.
Logistics and content exposure
Fully remote and asynchronous, with work assigned in project batches; hours are flexible and volume varies with customer demand rather than arriving as a steady full-time load. Observed pay is $20–22/hr — a band reported by contributors, not a guarantee. The work is text-only. Some projects touch sensitive material — bias, misinformation, harmful behaviors — and Mercor states that topics are communicated before exposure, that participation in higher-sensitivity projects is optional, and that guidelines and wellness resources are provided. Be honest with yourself about that before you apply.