What the work actually is
You write the problem, then you write the answer key. A task might start with a callback on a 4-ton split system: 12°F superheat, high subcooling, suction line sweating back to the compressor, homeowner says it runs constantly and never cools. You document the symptom set, the measurements a tech would take next and in what order, what each reading rules in or out, and the repair — plus the reasoning that connects the refrigeration cycle behavior to the conclusion. The reference answer matters as much as the scenario; a model can be graded only against reasoning that is explicit.
The other half is grading. You read model output against your key and score it for technical accuracy, diagnostic efficiency, and safety. The safety dimension is where trade expertise earns its rate: models produce advice that sounds like a seasoned tech and would flood a compressor with liquid, top off a system that has an active leak, or walk a homeowner past a cracked heat exchanger without mentioning combustion analysis. Catching plausible-but-lethal is the core skill.
What the screen looks for
- Verifiable credentials. EPA Section 608 (type and issuing organization), state license number where your state requires one, and where and when you did field work.
- Depth under follow-up. Expect the screener to push on a diagnosis you give — asking what reading would change your mind, or why you ruled out the obvious alternative. Confident answers that collapse on the second question score worse than cautious ones that hold.
- Written reasoning. Your writing sample is the deliverable. Chained, ordered explanation beats trade shorthand.
- Grading judgment. Can you separate "different approach than mine" from "wrong," and "inefficient" from "unsafe"?
Logistics
Fully remote and asynchronous — no tools, no travel, no site visits. You pick your hours; the platform asks for at least 10 hours a week so that batch turnarounds stay predictable. Pay is weekly via Stripe, at a rate set within the observed $50–90/hr band based on credentials, specialty depth (commercial refrigeration, BAS, geothermal), and calibration performance on early batches. Volume fluctuates with lab demand, and most contributors treat it as something that fits around a service schedule rather than as a replacement for one.