What the work actually involves
This is task-based evaluation rather than open-ended chat rating. You give an AI assistant a realistic instruction — reschedule a dinner with three people, book a haircut for next Tuesday afternoon, chase an unanswered thread about splitting a bill — and then you check what it actually did. That means opening the sent mail, reading the calendar invite it created, confirming whether the booking went through and at the right time zone, and noting every place the output diverged from the instruction. You then apply the project's grading rubric to score, rank, or annotate the result, and submit structured findings through micro1's tooling.
Coverage matters as much as scoring. Projects like this need edge cases: ambiguous instructions, multi-person coordination where one party never replies, invites that collide with existing events, recurring meetings, attachments, partial failures where the assistant did three of four steps. Evaluators who only run the happy path produce data nobody can learn from.
What the platform screens for
- Genuine daily use of Gmail and Google Calendar — not familiarity, but running your own life on them: accepting and declining invites, moving things, managing threads.
- Independent online booking habits — restaurants, appointments, travel, deliveries done by you, not by someone else on your behalf.
- Hands-on AI assistant use for real tasks, including the instinct to verify rather than assume the output is correct.
- Hard device and account gates: US-based, iPhone with iMessage, and an active Facebook or Instagram account. These are infrastructure for the task scenarios, not preferences.
- Written clarity in your annotations — a failure note that says "wrong time" is worthless next to one that says "set 3pm ET when the thread agreed on 3pm PT."
Logistics
Remote, contractor, US only. Work is asynchronous and milestone-driven: you pick up task batches, complete them under standardized conditions, and submit within the project's windows. Expect ongoing quality review — calibration against other contributors is normal, and consistent divergence from the rubric is how people get cycled off. Observed pay is $15–30/hr; banding typically reflects prior evaluation experience and calibration performance rather than negotiation. Hours are generally flexible, but batches often carry turnaround deadlines, so blocks of a few hours are more workable than fifteen minutes here and there.