What the work actually involves

You are not shipping product Java — you are producing training and evaluation material for large language models. On a typical day you might write a clean, idiomatic reference implementation for a backend task, compare two model attempts at the same problem and rank them, then write a rationale explaining precisely why one is better: a subtle concurrency bug, a resource leak from an unclosed stream, an API misuse that compiles but behaves wrong under load. The rationale is the deliverable as much as the code is. Other days lean toward supervised fine-tuning work — authoring task-specific prompt/response pairs — or peer-reviewing other contributors' submissions for correctness, annotation quality, and consistency with the project rubric.

Turing runs these engagements for AI labs, so task types and rubrics can shift between weeks. Expect specification documents, rubric updates, and reviewer feedback loops. Contributors who do well treat the rubric as the source of truth even when their personal engineering taste disagrees, and flag ambiguity rather than silently resolving it.

What the screen looks for

The evaluation is roughly 75 minutes: a 60-minute technical round plus a 15-minute cultural and offer conversation. The technical portion probes real Java depth — collections and their performance characteristics, concurrency, the memory model, exception design, testing practice, build tooling, and how you reason about modular, secure, maintainable backend architecture. Expect follow-ups that push past your first answer. A second axis is evaluation judgment: can you look at two plausible model answers and articulate a defensible ranking with specific technical evidence, not vibes. Written English matters here more than in most engineering roles, because your rationales become training signal.

Logistics

  • Fully remote, contractor assignment — no medical or paid leave.
  • Stated contract duration is one month, with renewal common but not promised; start dates are typically fast, often the following week.
  • Commitment tiers of 20, 30, or 40 hours per week, with a minimum of four hours per day and four hours of daily overlap with US Pacific time.
  • Pay is undisclosed on this listing. Turing rates for coding-domain work are usually set per project and hourly; ask for the specific band during the offer conversation rather than assuming a published figure applies.