What the work involves

Tasks sit at the intersection of machine learning and biological data analysis. Expect to write and document model code against real omics datasets — supervised and unsupervised approaches to gene expression, variant interpretation, pathway modelling, single-cell embeddings — and to construct or repair pipelines that integrate ML components reproducibly. Some assignments are model development from a specification; others are research-style tasks where the correct approach is itself part of the deliverable. Written documentation of methodology, assumptions, and failure modes is treated as part of the work product, not an afterthought, because your reasoning is often the artifact being collected.

What the platform screens for

AfterQuery's screen is AI-led and follow-up heavy. It looks for two things a generalist ML engineer usually cannot fake: fluency with the statistical peculiarities of biological data (batch effects, compositionality, n≪p, dropout in scRNA-seq, label noise in clinical annotations) and the ability to say plainly when a model result is an artifact rather than a finding. Your first-author publication will be asked about in specifics — what you did versus your co-authors, what the analysis pipeline actually was, what you would change now. Tooling questions (GATK, STAR, DESeq2, Nextflow, Snakemake, scVI, DNABERT, Enformer) are checked for hands-on use, not name recognition.

Logistics

  • Fully remote and asynchronous; no fixed meeting hours
  • Project-by-project assignment across a 2–3 week window, with 10–20 hours per week expected on projects you accept
  • Early-stage Y Combinator-backed client; scope can shift between projects
  • Pay observed at $150–200/hour, project-based and not guaranteed for any given assignment

This suits a postdoc, senior graduate researcher, or industry computational biologist who can work unsupervised, reads a task spec closely, and writes clearly enough that someone else can reproduce the result.