What the work involves
This is product engineering on top of model infrastructure, not model research. You will be shipping end-to-end features: a React or Next.js frontend, backend services in Python or Node.js/TypeScript, and the API layer between them and LLM inference endpoints. A large share of the day-to-day is handling the awkward parts of inference in a web app — token streaming over SSE or WebSockets, requests that run for minutes rather than milliseconds, retries and partial failures, job queues and webhooks for long-running tasks, and keeping the UI honest while all of that happens.
Deployment lives on GCP. Expect Cloud Run and Cloud Functions for services, Firestore or Cloud SQL for state, GCS for artifacts, and collaboration with infrastructure/DevOps counterparts on monitoring, reliability and performance. You will also design and maintain integrations against relational and NoSQL stores and third-party services, with Docker, Git and CI/CD as the normal working toolchain.
What the platform screens for
Turing's process is a profile review followed by an interview with the delivery team where warranted, so the screen leans on specifics you can name: which LLM APIs you integrated, how you handled streaming and timeouts, which GCP services you actually deployed to versus read about, and what broke in production. Generic fullstack claims without an inference-integration story are the usual reason a profile stalls. Being able to talk concretely about cost, latency and failure behaviour of model-backed endpoints separates candidates here.
Logistics
- Remote, contractor engagement — no medical or paid leave
- 8 hours per day, with at least 4 hours overlapping PST
- Initial contract duration of 2 months, immediate start expected
- Pay band undisclosed on this listing; Turing rates are typically negotiated per engagement and by region