What the work actually involves

You are not operating production infrastructure here — you are producing the material a model learns systems engineering from. Day to day that means authoring hard, realistic tasks in one or more of four areas: custom GPU kernel work (CUDA, Triton, Pallas), performance profiling and trace interpretation (torch.profiler, Kineto, Nsight, XLA/JAX profiler), debugging distributed or accelerator-bound workloads, and high-throughput LLM serving (vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, continuous batching). Each task ships with a reference solution you can defend line by line, plus written reasoning explaining why a given optimization or fix is correct and what the alternatives cost.

The other half is evaluation: reading model-generated or peer-written solutions and grading them against rubrics you help build. A rubric here has to encode real engineering judgment — that an occupancy improvement which regresses memory bandwidth is not a win, that a flame graph showing a gap is not the same as a diagnosis, that throughput gains bought with p99 latency may be the wrong trade for the stated workload. Your written feedback gets read by reviewers and other SMEs, so it has to hold up under challenge.

What the screen looks for

  • Hands-on depth, not familiarity. Two-plus years of professional ML systems, ML infra, serving, or accelerator performance work. Applied modelling and data science backgrounds are explicitly not the target profile.
  • At least one of the four pillars at real depth, with more than one a distinct advantage. Expect follow-ups that go a layer below whatever you claim — if you say you wrote Triton kernels, be ready to discuss block sizing, masking, and how you validated numerics.
  • Production PyTorch and/or JAX, with framework-level work (custom ops, FSDP/DDP/DeepSpeed/Megatron, compiler or graph-level changes) treated as a strong plus.
  • Written clarity. Much of the deliverable is prose explaining technical decisions, and the screen weighs how precisely you explain trade-offs.

Logistics

This is W-2 employment with Cincinnatus LLC, the employer of record, with placement into a leading AI lab's extended workforce. It is a 40-hour full-time weekday engagement with an explicit no-conflicts, no-other-engagements condition — this is not a nights-and-weekends side arrangement. Work is remote and largely asynchronous, with collaboration across other subject-matter experts to keep task and rubric standards consistent. The $90–120/hr band reflects rates observed on this listing and is not a guarantee; final rate depends on assessed depth and area coverage.