What the work actually is
You write customer support problems that are hard for a model in a specific way, then judge whether the model's answer would have survived contact with a real customer. A task might be a chat transcript where the customer is factually wrong about the warranty terms but emotionally justified; a refund request that sits just outside policy where the retention-versus-precedent tradeoff is the whole point; a three-channel escalation where the phone agent promised something the email record contradicts. You supply the scenario, the reference answer, and the reasoning trail that explains why one response is correct and a plausible-sounding alternative is not.
The second half of the job is evaluation. You read model outputs and mark them against your own rubric — did it acknowledge before solving, did it invent a policy that does not exist, did it apologise so much it admitted liability, did it hand off when it should have owned the issue. Feedback has to be specific enough for someone who has never worked a queue to act on. "Not empathetic enough" is rejected; "opened with the policy citation before acknowledging the failed delivery, which is the sequencing that generates repeat contacts" is the standard.
What the screen looks for
- Real queue time, not adjacent experience. Team leads, QA analysts and CX designers do well, but only if they can still speak concretely about handling contacts themselves.
- Channel breadth. Phone de-escalation, async email tone, chat concurrency and social escalation all have different failure modes. Candidates who have only worked one channel are screened for depth instead.
- Ability to articulate a rubric. Most applicants can say a response is bad. Fewer can say why in terms another grader would reproduce.
- Writing under scrutiny. Your scenarios become training data; ambiguity in your prompt becomes noise in the model.
Logistics
Fully remote and asynchronous — no scheduled shifts, no live customer contact. Contributors typically commit 10–20 hours a week and pick up batched task sets with turnaround windows measured in days rather than hours. Pay is hourly within the observed $35–75 range; the upper end has been associated with QA, training, or multi-industry escalation backgrounds rather than tenure alone. Engagement is contractor-based and project-dependent, with volume varying between batches.