What the work actually involves
You will spend most of your time producing reference-quality technical artifacts and then judging the model's attempts at the same. A typical task queue mixes authoring (write an architecture decision record for a service migration, build a data quality validation test plan as a spreadsheet, draft a developer onboarding guide) with evaluation (compare two model-generated API integration guides and write a defensible rationale for which is better and why). The document is the deliverable — formatting, table structure, slide hierarchy, and cross-references count as much as the technical content, because the lab is training the model on professional artifact craft, not just prose.
Evaluation work demands written justification that another engineer could audit. Saying a pipeline doc is "incomplete" earns nothing; identifying that it omits backfill behavior, says nothing about late-arriving records, and invents a Spark config flag that does not exist is the work. Hallucinated API endpoints, silently wrong SQL, unstated schema assumptions, and version drift are the failure classes you are being paid to catch.
What the screen looks for
- Four-plus years in a real engineering or data role — shipped systems, not coursework or adjacent PM work.
- Artifacts you personally wrote: design docs, ADRs, runbooks, API specifications, ETL documentation, readiness reviews.
- Prior evaluation judgment — code review, design review, model output assessment, or system readiness sign-off where you documented tradeoffs for a mixed audience.
- Tooling fluency in documents, spreadsheets, and slides at a level where you can articulate why a particular table or diagram structure serves the reader.
The screen is an AI voice interview. It asks follow-ups, and shallow answers get probed rather than accepted. Concrete systems, numbers, and named tradeoffs survive that; rehearsed summaries do not.
Logistics
Fully remote and asynchronous — no standing meetings, no fixed hours, no timezone requirement stated. Commitment is flexible between 5 and 20 hours per week, with more available if you want it. Work is task-based, so throughput and quality reviews determine how much flows to you after onboarding.