What the work actually is

You are not testing hardware here; you are writing the exam. Each task starts from a scenario you have genuinely lived — a load cell reading that drifted after a temperature soak, a vibration fixture with a suspect accelerometer channel, a production check that failed CpK for reasons that weren't the part — and you build it out into a self-contained problem an AI model must work through. That means authoring the setup, producing or synthesizing the artifacts (test logs, calibration records, raw sensor exports, NCR or failure reports, usually as CSVs or spreadsheets), defining the correct reasoning path including how the data should be read and what the next step in the lab would be, and then writing a grading rubric of 35 or more discrete, checkable items covering instrumentation logic, data interpretation, and diagnosis.

The rubric is where most of the effort lands. Items must be atomic and gradeable by someone other than you: "identifies that the channel gain was set for a ±5 V sensor on a ±10 V output" scores; "understands signal conditioning" does not. Reviewers reject tasks where the rubric can be satisfied by a plausible-sounding answer that a real technician would recognize as wrong, and where the scenario is a textbook exercise dressed up — clean numbers, one variable, no ambiguity.

What the screen looks for

  • Verifiable hands-on time. Four-plus years in a lab, test cell, metrology room, or production test environment. Which rigs, which DAQ, which sensors, what you calibrated and to what standard.
  • Failure-analysis judgment. Whether you can distinguish an instrumentation artifact from a real part failure, and explain how you ruled the other out.
  • Documentation discipline. Evidence you have written test procedures, reports, or NCRs that someone else acted on.
  • Rubric thinking. Can you decompose a diagnosis into scoreable steps and name the wrong-but-tempting answers?

No AI or machine-learning background is expected or asked for.

Logistics

Fully remote and asynchronous — you choose your hours. Pay is per accepted task, not hourly; the observed band for this kind of work is roughly $30–70/hr equivalent, which depends heavily on how fast you assemble datasets and rubrics, and is not guaranteed. A weekly minimum submission count applies. Hiring moves fast: roles are typically filled within 48 hours and first tasks are expected within a day or two of onboarding, so apply when you actually have bandwidth.