What the work actually involves
You will spend most of your time probing models in the areas of finance you know best and documenting where they break. That means building prompts drawn from real work — a revenue build with a channel mix change, a comparable-companies screen, an accretion/dilution question, a lease-accounting treatment, a risk-limit breach scenario — then judging the model's answer against what a competent practitioner would produce. Alongside evaluation, you write rubrics: explicit criteria that let a researcher or another expert score the same output and reach the same verdict. Some tasks involve reference answers with worked steps; others ask you to compare two model responses and defend the ranking in writing.
What the screen looks for
Turing's process is an AI-led interview followed by domain tasks. The screen is checking that your stated experience survives follow-up questions: which desk, which deal types, what you personally built versus what your team shipped. It also tests evaluation judgment — whether you can separate a confident, well-formatted answer from a correct one, and whether you can name the specific error (wrong discount convention, double-counted synergy, misapplied revenue recognition) rather than calling an answer "weak." Written English matters more than usual here, because the rubric text is the deliverable.
Logistics
- Fully remote, US-based, largely asynchronous with occasional sync calls with researchers
- 10–30 hrs/week, self-scheduled; roughly one month with extension possible based on performance
- Observed rate around $100/hr, varying with domain and depth of experience — not guaranteed
- No prior AI or machine learning experience required; CFA, CPA/CA, or a finance MBA is a bonus, not a gate