Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Sr. SWE Datacenter Automation

Home > Javascript/Typescript programming jobs

Sr. SWE Datacenter Automation in USA

  • Zipline
  • Full time
  • Email
  • South San Francisco, CA

Responsibilities

  • Own datacenter compute and storage lifecycle, including bare-metal provisioning, hypervisor management, SAN/NVMe storage clusters, network configuration, and Kubernetes cluster lifecycle.
  • Build automation for PXE/firmware workflows, dynamic inventory, image generation, configuration drift detection, and automated recovery.
  • Define and deliver reliability targets for provisioning, node commissioning, cluster upgrades, and mean time to recover.
  • Lead infrastructure incident response, on-call rotations, incident command, postmortems, and follow-up actions.
  • Maintain monitoring, alerting, and dashboards for hardware, hypervisors, storage, Kubernetes control planes, and autoscaling.
  • Manage infrastructure cost, capacity, reclamation, firmware and hypervisor patching, cold standby, and failover procedures.
  • Perform datacenter racking, cabling, hardware troubleshooting, forensic log capture, and coordination of physical repairs.

Requirements

  • At least 5 years of engineering experience, including at least 4 years owning production datacenter, virtualization, or infrastructure automation systems.
  • Hands-on expertise with bare-metal provisioning and imaging using PXE/iPXE, IPMI, and Redfish.
  • Experience operating KVM/QEMU or ESXi hypervisors and Ceph, NVMeoF, or SAN storage systems at scale.
  • Production Kubernetes operations experience covering provisioning, upgrades, highly available control planes, cluster APIs, CNI and CSI troubleshooting, and multi-cluster workload scheduling.
  • Production automation and coding skills in Python, Go, or Rust, plus experience with Terraform, Ansible, Helm, and CI/CD pipelines.
  • Strong networking fundamentals including VLANs, BGP/EVPN leaf-spine networks, LACP, routing, cluster networking, and storage-fabric troubleshooting.
  • Experience with on-call operations, postmortems, SLIs/SLOs, and reliability improvements.
  • Ability to work hands-on in datacenters and operate in regulated or safety-sensitive environments with change-control and audit processes.
  • Ability to work regularly onsite in South San Francisco, travel occasionally to datacenter or field locations, and collaborate across technical and operations teams.

Benefits

  • South San Francisco-based role with regular onsite presence and in-office cadence.
  • Participation in on-call rotations and occasional travel to partner datacenters or field sites.
  • Equal opportunity workplace committed to diversity and inclusion.

Zipline

Zipline is redefining logistics with its instant delivery system that addresses urgent access challenges for a diverse range of customers, including governments and businesses. By leveraging advanced technology like robotics and autonomy, Zipline ensures equitable access to essential goods, wheth...

Similar positions

Sr. Software Engineer B2

  • Cognizant
  • Full time
  • Others
  • 09/13/2026
  • Bengaluru, India

Sr. Fullstack Developer

  • iSpace, Inc.
  • Contract
  • USA
  • 09/12/2026
  • Remote

Sr. IT Full Stack Developer

  • Databricks
  • Full time
  • Others
  • 09/11/2026
  • Bengaluru, India

Platform Enablement & Automation Engineer

  • Payoneer
  • Full time
  • Others
  • 09/11/2026
  • Hong Kong, Hong Kong

Sr. Fullstack Product Engineer

  • Salt Ai
  • Full time
  • USA
  • 09/08/2026
  • Salary: $140k-$180k
  • Remote