What the work involves
Each task starts with a user goal drawn from ordinary small business life: chasing an unpaid invoice, drafting a supplier email, reading a P&L, pricing a new service, writing a social post for a slow month. You turn that goal into a realistic prompt — often attaching one of your own business documents (a spreadsheet, a PDF quote, a photo of a receipt or shelf display) — then run the same scenario against multiple chatbots for up to five turns each. Afterwards you write a structured comparative evaluation: which response was actually usable, where a model invented a figure or misread a column, which one produced advice that sounds professional but would get an owner into trouble with a customer or a tax authority.
Deliverables are the conversation transcripts plus the evaluation form. The judgement being paid for is operator judgement, not writing polish — whether the AI's answer would survive contact with a real customer, a real cash-flow constraint, or a real deadline.
What the platform screens for
- Genuine operating experience. Expect follow-ups asking what you actually do, how many people you employ or contract, what your invoicing and bookkeeping look like. Vague "I've consulted for SMEs" answers get probed hard.
- Document access. The project is built around real input files in English. Screeners will ask what kinds of documents you can supply and whether you can redact client and personal data before uploading.
- Evaluation discipline. Can you separate "I liked this tone" from "this response contained a factual error"? Can you follow a rubric you disagree with, and flag the disagreement separately rather than scoring around it?
- Comfort with the tooling. Multi-chatbot interfaces, file uploads, transcript capture.
Logistics
Fully remote, asynchronous, project-based with a defined task count over a stated 16-week duration. Work is self-scheduled; most contributors fit tasks around running their business rather than committing to fixed hours. Pay is undisclosed on this listing — Turing typically quotes a per-task or hourly rate at the offer stage, and nothing here should be treated as a guaranteed band. Confirm the rate, the task count, and how rejected tasks are handled before you start.