What the work actually involves

You pick problems from the territory you already know cold — a stripped ARM64 binary with a custom packer, a protocol that reuses a nonce, a homegrown block cipher with a weak key schedule, an obfuscated loader with anti-debug checks — and turn them into evaluation items. Each item needs a scenario a competent practitioner would recognize as realistic, a reference solution written out at expert depth, and a rubric that distinguishes a correct answer from one that merely sounds correct. Then you score model attempts against that rubric and write structured feedback explaining exactly where the reasoning held and where it broke.

The hard part is rarely the security. It's the rubric. Models are fluent in the vocabulary of both fields — they will name padding-oracle attacks, describe control-flow flattening, cite Coppersmith — while producing an attack that never actually recovers the key or a devirtualization narrative that doesn't match the bytes. Your job is to build grading criteria that catch that gap reproducibly, so a second reviewer applying your rubric reaches the same verdict.

What the platform screens for

  • Hands-on recency. Screeners probe for work you did, not work you read about. Expect follow-ups on tooling versions, specific target architectures, and what went wrong.
  • Depth under pressure. A single answer rarely settles a question; the screen drills two or three layers into whatever you claim.
  • Written precision. Reference solutions and feedback are the deliverable. Vague prose fails here even when the underlying analysis is right.
  • Calibration. Can you separate a confidently wrong model answer from an unconventional but valid one?

Logistics

Fully remote and asynchronous. Contributors typically work in blocks of a few hours, choosing their own times, with output measured in completed items and reviews rather than logged presence. Most people run this alongside a primary role. Expect an onboarding calibration round where your gradings are compared against other reviewers before you're given independent volume. Pay is hourly within the observed $100–170 band, set by experience and track depth — not guaranteed, and the platform sets the final figure.