What the work actually involves

You write short multi-turn conversations in Arabic — usually one to five turns — that force the model to reach for something it could only know from your own data: a thread in your Gmail, something you searched last month, a channel you watch, an earlier Gemini conversation. Then you judge what comes back. Was the claim about you actually supported by a data source, or was it a plausible-sounding inference? Was the personal detail woven in, or did the model announce it ("Since I see from your Gmail that…")? You stack-rank two responses side by side, write a rationale that points to the specific turn where the problem appeared, and pull the model's Debug Info to confirm which sources and chat summaries were really used. You also delete your evaluation conversations afterwards so they don't contaminate your own future chat history.

What the screen is looking for

Turing's process here is a Job Interest Form, a profile review, then a timed assessment that must be returned within 24 hours. The assessment is the real gate: it tests whether you can tell the difference between good personalization and confident guessing, and whether your written rationale in Arabic is specific enough that a reviewer who wasn't there could agree or disagree with it. Vague verdicts — "Response A felt more natural" — are the most common way people wash out. Expect probing on:

  • Arabic reading and writing at a level where you can critique register, naturalness and overnarrating, not just comprehension
  • Willingness to enable personal data sources on your real Google account, with informed understanding of what that means
  • Whether your prompt designs genuinely stress the feature rather than being generic questions any model could answer
  • Prior annotation, evaluation, moderation or analytical work, which is strongly preferred rather than strictly required

Logistics

Contractor engagement, roughly 3 months, remote. Minimum 4 hours per day and up to 40 hours per week, with 4 hours of overlap with PST — the team is staffed as a 24-hour global operation, so your local schedule matters less than that overlap window. Desktop or laptop and a stable connection required; phone-only work isn't viable given the side-by-side interface and Debug Info extraction. The rate observed on this listing is $15 per hour and is not guaranteed to hold across future cohorts or projects.