The work

You would build, operate, and improve highly available production infrastructure, primarily using Kubernetes and AWS. Day-to-day work includes investigating incidents, completing root-cause analysis, improving capacity and deployment reliability, and automating repetitive operational tasks with infrastructure and platform tooling.

Reliability and tooling

The role covers monitoring, alerting, logging, and tracing through Datadog or comparable observability platforms such as Prometheus or Grafana. You may also maintain Infrastructure as Code with Terraform, Pulumi, or equivalent technology, support CI/CD systems, and collaborate with software engineers on deployments and production architecture. Proficiency in at least one programming or scripting language is expected, with Python, Go, or Bash listed as examples.

What the platform screens for

Expect the screening to test whether you can describe specific production systems, their scale and failure modes, and the decisions you personally owned. Strong answers will connect technical changes to measurable improvements in SLOs, SLIs, observability, deployment reliability, infrastructure performance, capacity, or operational efficiency.

Practical details

This is a full-time engagement, and candidates must be able to make a full-time commitment. The listing does not specify whether the work is remote or asynchronous, nor does it provide a daily schedule or time-zone requirement.

Pay band: $200/hr