Establish service-quality strategies by defining service-level indicators and error budgets.
Operate, optimize, and improve large-scale production systems through maintainable code, reliability practices, on-call participation, and incident response.
Improve customer and internal user experience through performance engineering, observability adoption, and automation.
Apply technical judgment to ambiguous problems, prototypes, hardening decisions, and tradeoffs.
Raise the team’s technical bar through design feedback, code review, mentoring, and clear written communication.
Use AI-powered development tools across planning, implementation, review, testing, and iteration while maintaining independent judgment.
Requirements
Experience operating and optimizing large-scale distributed systems in production, including observability and self-healing techniques.
Expertise managing production services in cloud environments such as AWS, Azure, or GCP.
Experience with monitoring and observability tools such as Datadog, Grafana, Prometheus, and OpenTelemetry.
Experience operating production systems in a high-growth startup environment.
Familiarity with declarative production-infrastructure management and modern Infrastructure as Code tools such as Kubernetes and Terraform.
Benefits
Remote-first team with a flexible-first culture.
Equity stock options.
401(k) with a 5% company match and immediate vesting.