What the work actually involves

You are not maintaining someone's warehouse on-call. You are producing self-contained data engineering artifacts — a DBT model with tests, a Spark job that handles a skewed join, a Kafka consumer with a poison-message path, a Snowflake or Redshift schema with the loading logic behind it — along with the reasoning that explains why it is built that way. Much of the value sits in the documentation: data flow diagrams, architectural rationale, and clear write-ups of failure modes. A recurring task type is debugging: you are handed a pipeline that fails, times out, or silently drops rows, and you must identify the cause, fix it, and describe the diagnostic path you took.

What the platform screens for

The screen is AI-led and follow-up heavy. It probes whether your stated tool experience is real production experience or tutorial familiarity — expect to be asked about specific configuration, cost or partitioning decisions, and incidents you personally handled. Vague answers get drilled. The screener also tests written explanation, since your output is read by people who were not in the room: if you cannot describe a backfill strategy or an incremental model's edge cases in plain prose, the work product is weak regardless of the code. Bachelor's degree in CS, engineering, information systems, or a related technical field is stated as required, alongside 1–3 years of professional experience.

Logistics and pay

  • Fully remote and asynchronous; no fixed meeting hours
  • Project-by-project assignment across a 2–3 week engagement window, roughly 10–20 hours weekly on accepted projects
  • Observed band of $60–100/hr, described by the platform as starting at an estimated $60 and rising with task complexity and reviewed quality — rates are as observed, not guaranteed
  • Client is an early-stage Y Combinator-backed company; volume can be uneven, and accepting a project means committing to its deadline