What the work actually involves
Each task starts with you, not the model. You pick a real small business goal — reconciling an invoice discrepancy, drafting a supplier follow-up, reading a quarterly sales sheet, deciding whether a quote is worth accepting — and write a prompt that reflects how that situation actually arrives in a business owner's day. You then attach or reference a real Spanish-language business document (spreadsheet, PDF, image of a receipt or contract) and run the same prompt through multiple chatbots, capped at five turns per conversation. You score each response on clarity, usefulness and accuracy, write a comparative judgment saying which model served the business goal better and why, and submit the transcripts alongside your evaluation.
The document requirement is not decorative. Turing is testing multilingual, multimodal, document-grounded behaviour, so the value of your task depends on the model having to read something real in Spanish rather than answer a generic question. Expect topics to span marketing content, day-to-day operations and customer handling, informal market research, and financial planning off spreadsheet data.
What the screen looks for
- Evidence you have actually run or closely operated a small business — revenue scale, headcount, what you personally handled versus delegated.
- Confirmation that you hold usable Spanish-language business documents and are willing to use them, including how you would redact identifying client or personal data.
- Whether you can separate a confident-sounding answer from a correct one. Business advice is easy to write and hard to verify; the screen probes whether you check numbers against the source document rather than rewarding fluency.
- Comfort following a rubric you did not write, including scoring a response you personally disagree with when the guideline says it satisfies the criterion.
Logistics
Fully remote and asynchronous, with no fixed hours. Work is allocated as a defined batch of evaluation tasks across a stated 16-week window, so throughput varies week to week rather than arriving as a steady shift. Pay is not disclosed in the posting; ask about per-task rate, expected time per task and batch volume before committing, since project-based AI evaluation work is typically compensated per completed and accepted task.