What the work actually is

You are not shipping production code. You are deciding what "correct" and "excellent" mean for a model attempting a systems-engineering problem you could solve yourself. In practice a task cycle looks like: pick a realistic scenario from your own domain — a KV-cache eviction policy that regresses throughput under long-context load, a CUDA kernel with a shared-memory bank conflict, a tensor-parallel sharding choice that silently changes numerics — specify it precisely enough that a strong engineer could attempt it cold, write a reference solution, then grade model outputs against a rubric you helped define. Written justification matters as much as the verdict; the trace of your reasoning is part of the deliverable.

What the screen is looking for

  • Hands-on depth in one narrow area, not breadth. Real experience with vLLM, SGLang, TensorRT-LLM internals, custom CUDA/Triton kernels, quantization schemes, or optimizer implementation. Screens push on specifics: which version, what did you measure, what broke.
  • Systems ownership at a recognized technology or AI company — a role where you owned latency, throughput, or memory budgets rather than consuming someone else's serving stack.
  • Grading judgment. Whether you can separate a plausible-sounding wrong answer from a correct-but-unfamiliar one, and whether you'd notice a model that produces the right number for the wrong reason.
  • Writing. Technical explanation that a reviewer who doesn't share your specialty can follow.

Logistics

Fully remote, contract, asynchronous. Part-time by design — most contributors run this alongside a full-time engineering job, taking batches of tasks when they have hours. Expect variability in volume: work arrives in project waves tied to specific model evaluations, and some waves target narrow specialties (SGLang, Mamba/Mamba2, Granite-family models) more than others. Pay in the $100–150/hr band has been observed for this category and typically tracks specialization depth rather than tenure.