Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Senior Infrastructure Engineer, SRE

Home > Python programming jobs

Senior Infrastructure Engineer, SRE in USA

  • Rocket Money
  • Full time
  • Email
  • Remote

Responsibilities

  • Build and improve the reliability and resiliency of production systems and services.
  • Establish and regularly review SLIs, SLOs, and error budgets for critical services and user journeys.
  • Own disaster recovery strategy, including recovery objectives, failover and restore paths, and regular exercises.
  • Partner with product engineering teams to improve service ownership, instrumentation, and user-focused metrics.
  • Evolve observability standards for metrics, tracing, logs, instrumentation, alert quality, and cost.
  • Strengthen incident response by tuning paging thresholds, maintaining runbooks, and tracking postmortem actions.
  • Contribute to cloud infrastructure build-outs, platform backlog work, automation, and a shared on-call rotation of one week every six weeks.
  • Lead or contribute to reliability and observability modernization, internal tooling, game days, chaos experiments, and disaster recovery exercises.

Requirements

  • 5+ years of hands-on cloud or infrastructure engineering experience, including substantial reliability and production operations work at scale.
  • Experience defining SLIs and SLOs for real production services and evaluating their operational impact.
  • Hands-on production experience with an observability platform; Datadog is strongly preferred.
  • Ability to write code in Python, Go, TypeScript, or a similar language for tooling, debugging, and automation.
  • Production Terraform experience and comfort working in AWS.
  • Experience building or operating disaster recovery plans, including recovery goals, failover and restore procedures, and drills.
  • Experience being on-call for services and improving alert quality.
  • Preference for enabling teams with paved roads and good defaults rather than mandates.
  • Preferred experience leading reliability or observability modernization projects and delivering their implementations.
  • Preferred experience building internal tooling, libraries, or instrumentation standards, running game days or chaos experiments, and reducing observability costs.

Benefits

  • Health, dental, and vision plans.
  • 401k matching.
  • Unlimited PTO.
  • Competitive pay, plus bonus and benefits.
  • Daily lunch, snacks, and coffee for in-office employees.
  • Commuter benefits for in-office employees.
  • Shared on-call rotation of one week every six weeks.

Salary: $150k - $185k/yr

Rocket Money

With more than 3.4 million members, Rocket Money helps users manage subscriptions, lower recurring bills, create budgets, track spending, and automate savings toward their financial goals. By combining financial insights with practical tools, the platform enables users to make smarter decisions a...

Similar positions

Senior Devops Engineer

  • Woliba
  • Full time
  • USA
  • 09/13/2026
  • Remote

Senior Core Infrastructure Engineer

  • Oracle
  • Full time
  • USA
  • 09/13/2026
  • Salary: Competitive
  • Nashville, TN

Senior Software Engineer, ML Ops

  • PathAI
  • Full time
  • USA
  • 09/13/2026
  • Salary: $128k - $196k/yr
  • Boston, MA

Senior Backend Software Engineer

  • LangChain
  • Full time
  • USA
  • 09/13/2026
  • Salary: $180k - $240k/yr
  • Boston, MA

Senior/Staff Full Stack Engineer

  • Workstream
  • Full time
  • Canada
  • 09/13/2026
  • Salary: Competitive
  • Vancouver