What the work actually is

You write the underwriting file the model should have produced, then judge whether it did. A typical task starts with a scenario you or another contributor authored — a mid-size habitational risk with prior water losses, a life application with an abnormal stress echo, a marine cargo submission with a questionable warranty — and asks for a work product: an accept/decline/refer decision, a rating and pricing rationale, a coverage analysis against specific form language, or a referral write-up to a senior authority. That becomes the "golden" reference. On grading passes, you score AI-generated responses against rubrics covering risk-selection accuracy, provision interpretation, pricing correctness, and whether the documented rationale actually supports the decision. The written feedback matters as much as the score: researchers use it to diagnose why a model misread an exclusion or fabricated a rating factor.

The failure modes you are hunting are specific. Models confidently cite endorsements that do not exist, apply the wrong trigger to an occurrence versus claims-made form, treat a condition as an exclusion, botch arithmetic on layered limits or experience mods, and — most commonly — reach a defensible conclusion with a rationale that would not survive an audit or a regulatory file review. You flag all of it in writing.

What the screen is looking for

Mercor's screening is AI-led and pushes on depth. Expect to name your lines, your authority limits, and the systems and forms you worked in, then get follow-ups that only someone who actually underwrote would answer cleanly — how you priced a specific class, what you referred upward and why, how you documented a decline so it held up. Credentials like CPCU, CLU, AINS, ARM, AU, or FALU are a plus, not a gate; two years of real underwriting authority is the floor. The other thing being tested is boundary discipline: an underwriting decision is not a producer's recommendation and is not a legal coverage opinion, and candidates who blur those lines score poorly.

Logistics

  • Fully remote and largely asynchronous, with task batches you claim and complete on your own schedule
  • Minimum 20 hours per week; 40+ preferred, and throughput tends to determine how much work routes to you
  • Live onboarding office hours and periodic calibration sessions — these are synchronous and expected
  • Rolling review, immediate start; $80/hr is the observed rate for this listing and varies by line, seniority, and task type