Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Senior/Staff Kubernetes Infrastructure Engineer

Home > Python programming jobs

Senior/Staff Kubernetes Infrastructure Engineer in Remote

  • Worldwide

Responsibilities

  • Design, automate, validate, and deliver the complete lifecycle of customer compute environments from provisioning through upgrades, recovery, and decommissioning.
  • Use AI to automate and accelerate infrastructure delivery and operations.
  • Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads.
  • Build and maintain Linux images and automated operating-system provisioning workflows.
  • Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring.
  • Design Kubernetes and data-center networking using Cilium, Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP.
  • Configure distributed and shared storage for high-performance workloads.
  • Build monitoring, alerting, diagnostics, and automated recovery for customer environments.
  • Develop reusable tooling, standards, documentation, and runbooks.
  • Collaborate with customers and internal teams to translate workload requirements into infrastructure designs.

Requirements

  • At least 5 years of experience building and operating production Linux infrastructure.
  • Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, highly available control planes, etcd, containerd, CNI, CSI, ingress, load balancing, observability, security, and troubleshooting.
  • Experience with Linux virtualization using KVM/QEMU, libvirt, and VFIO device passthrough.
  • Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry.
  • Strong networking fundamentals covering TCP/IP, Layer 2/Layer 3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark.
  • Practical scripting experience and experience with configuration-management tools such as Ansible.
  • Ability to diagnose complex cross-layer infrastructure issues and drive technical decisions across teams.
  • Production Slurm experience is preferred.
  • Preferred experience includes NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, IMEX, hugepages, NUMA, CPU pinning, SR-IOV, and DPDK.
  • Preferred distributed-storage experience includes Ceph, Lustre, or Weka.
  • Preferred experience includes KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, and Nornir.
  • AI training, inference, or distributed GPU workload infrastructure experience is preferred.
  • Python or Go proficiency is preferred.

Salary: $180k - $250k/yr

fal

fal is a generative media platform that provides developers with access to the world's best generative image, video, and audio models through a unified API. Trusted by over 2.5 million developers and leading companies, fal offers the fastest inference engine for diffusion models, on-demand server...

Similar positions

Backend Engineer

  • Lendbuzz
  • Full time
  • Others
  • 10/02/2026
  • Salary: Competitive
  • Tel Aviv-Yafo, Israel

Product Security Engineer

  • Bugcrowd
  • Full time
  • Remote
  • 10/02/2026
  • Costa Rica

Senior Data Analytics Engineer

  • MariaDB plc
  • Full time
  • Remote
  • 10/02/2026
  • Salary: Competitive
  • India

Staff Engineer SDET

  • Aviatrix
  • Full time
  • USA
  • 10/02/2026
  • Salary: $189k - $223k/yr
  • Santa Clara, CA

Security Engineer

  • Deepgram
  • Full time
  • USA
  • 10/02/2026
  • Salary: $150k - $245k/yr
  • Remote