The work

You are a task designer for a scientific reasoning benchmark, not an annotator. Each task starts from a real single-cell or spatial workflow you have actually run — QC and normalization choices, cell-type annotation, pseudotime ordering, RNA velocity, multi-omic integration, spatially variable gene detection, persistence-based analysis — and gets converted into a problem with a verifiable answer. Two flavors dominate: fully specified problems where the model must execute a long multi-step pipeline correctly, and partial-information problems where the model must plan a sequence of queries or in-silico experiments to recover something the data does not show directly.

After designing, you test. Every problem runs against current frontier models, and you tune it until it sits in the target band — hard enough that pattern matching fails, clean enough that a competent postdoc with the same tools would agree on the answer. Expect to throw work away when a model solves it on the first try, or when your own oracle turns out to be ambiguous under a second reading.

What the screen looks for

  • Real tool mileage. Which scanpy defaults you have overridden and why; where scvelo's dynamical model misbehaves; what squidpy's neighborhood tests assume; how you set gudhi filtration parameters on gene expression data.
  • Verifiable provenance. Publications, preprints, open-source commits, or professional pipelines you can point to and discuss under follow-up.
  • Python engineering. You will write setup scripts, deterministic oracle functions, and validators that tolerate legitimate numerical variation without accepting wrong answers.
  • Problem-design instinct. The ability to distinguish genuinely hard from merely tedious, and to spot the shortcut that lets a model guess your answer.

Logistics

Fully remote and asynchronous, with work done in Linux remote compute sandboxes; containerized, reproducible setups are a plus. The listing asks for 15–20 hours per week minimum, and throughput matters more than fixed hours — this is described as the platform's highest-volume domain, so accepted contributors often see steady task flow. Observed pay for this listing is $70–100/hr, set by the platform based on screening outcome and domain depth; it is not a guarantee.