Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Senior Staff Service Reliability and

Home > Python programming jobs

Senior Staff Service Reliability and in USA new

  • Santa Clara, CA

Responsibilities

  • Own the technical strategy and multi-year roadmap for operational excellence and production readiness across development, pre-production, and production.
  • Define and govern the New Service Introduction framework and production-readiness reviews covering architecture, security, resilience, capacity, observability, supportability, and releases.
  • Establish service ownership standards for catalogs, accountable owners, dependency maps, runbooks, support models, escalation paths, recovery objectives, and on-call readiness.
  • Lead the architecture and evolution of the shared observability platform, including standards for logs, metrics, distributed traces, profiles, dashboards, alerts, synthetic monitoring, telemetry quality, retention, sampling, cardinality, and cost controls.
  • Own reliability governance involving SLIs, SLOs, error budgets, service health, customer impact, and escalation mechanisms.
  • Advance incident-management maturity through severity classification, incident command, communications, automated evidence collection, post-incident reviews, remediation tracking, and systemic fixes.
  • Lead capacity and efficiency management across demand forecasting, cloud and Kubernetes capacity, performance testing, scaling thresholds, headroom, rightsizing, and capacity-risk reviews.
  • Design and govern AIOps and secure AI-agent workflows for event correlation, alert-noise reduction, predictive detection, root-cause analysis, autonomous triage, remediation, and controlled self-healing.
  • Improve on-call effectiveness through rotation design, operational-readiness standards, escalation policies, diagnostic automation, alert-quality management, and follow-the-sun handoffs.
  • Provide hands-on technical leadership during major incidents, reliability investigations, architectural reviews, resilience exercises, and critical service launches.
  • Use operational, incident, service-level, capacity, change, and automation data to prioritize continuous improvement.

Requirements

  • 12+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
  • Recent experience designing and operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Deep understanding of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes.
  • Experience owning observability architecture and governing metrics, logs, traces, SLIs, SLOs, and error budgets.
  • Experience establishing reliability and operational-readiness standards for business-critical services.
  • Hands-on experience with failure experiments, disaster-recovery exercises, and validated service failovers.
  • Experience commanding SEV1 or SEV2 incidents, coordinating technical and executive communications, and driving systemic remediation.
  • Experience delivering measurable reliability outcomes involving availability, latency, MTTR, change-failure rate, alert quality, and error-budget adherence.
  • Experience with capacity forecasting, performance testing, scaling strategies, and cloud and Kubernetes resource management.
  • Strong software engineering and automation skills using Python or Go, infrastructure as code, and modern delivery toolchains.
  • Evidence of multi-team technical leadership through architecture reviews, standards, coaching, and mechanisms adopted beyond one team.
  • Ability to influence cross-functional stakeholders and deliver complex initiatives without direct management authority.
  • Preferred experience with risk prioritization, AIOps, autonomous remediation, AI-agent integrations, FinOps, load balancing, global traffic management, progressive delivery, and global follow-the-sun on-call models.
  • U.S. Person status, required authorization, or an applicable license exception may be necessary to access export-controlled technology and perform certain government-contract work.

Benefits

  • Comprehensive medical, dental, and vision plans.
  • Matching 401(k), unlimited PTO, paid holidays, parental/adoption leave, legal insurance, and a home technology stipend.
  • Located in Santa Clara, California, with up to 25% travel.
  • Equal opportunity workplace with accessibility and inclusion commitments.

Salary: $188k - $270k/yr

IonQ

IonQ, Inc. [NYSE: IONQ] is the world’s leading quantum platform and foundry - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ’s newest generation of quantum computers, the forthcoming IonQ Tempo, will be the latest in a line of cutting-edge system...

Similar positions

Senior Software Engineer

  • YipitData
  • Full time
  • USA
  • 08/17/2026
  • Salary: Competitive
  • Remote

Senior Software Engineer, Global Payroll

  • Rippling
  • Full time
  • Australia
  • 08/17/2026
  • Sydney

Staff Engineer - Recommendations

  • VRChat
  • Full time
  • Remote
  • 08/16/2026
  • USA, Canada

Staff+ Software Engineer, Safeguards Data

  • Anthropic
  • Full time
  • USA
  • 08/16/2026
  • Salary: $320,000 – $485,000
  • New York

Software Engineer, Fulfillment Core Services

  • Lyft
  • Full time
  • USA
  • 08/16/2026
  • Salary: $128,000 – $160,000
  • Seattle