Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Principal Infrastructure Engineer

Home > Golang programming jobs

Principal Infrastructure Engineer in Remote new

  • Mexico

Responsibilities

  • Own the technical architecture and evolution of core infrastructure as traffic, data volume, and workload complexity grow.
  • Build capacity models, run load and stress tests, diagnose bottlenecks, and improve throughput, latency, reliability, and cost efficiency.
  • Design resilient AWS account, IAM, networking, multi-AZ, and multi-region architectures.
  • Build and operate Kubernetes platforms, including cluster lifecycle, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
  • Scale and optimize Aurora RDS for MySQL and Postgres through query and index tuning, connection and replication improvements, failover planning, and safe migrations.
  • Implement reliability engineering practices including service-level objectives, error budgets, failure isolation, backpressure, load shedding, and safe retries.
  • Participate in on-call rotations, recover systems during major incidents, and implement corrective actions from postmortems.
  • Design, test, and document disaster recovery, backup, restore, and failover mechanisms against recovery objectives.
  • Build infrastructure-as-code, deployment, provisioning, upgrade, and operational automation to reduce manual toil.
  • Improve observability through metrics, logs, traces, dashboards, alerts, and visibility into customer impact.
  • Deliver safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback paths.
  • Build and evaluate AI-assisted tooling for incident investigation, capacity analysis, runbook automation, anomaly analysis, and toil reduction.
  • Write architecture proposals, prototype and benchmark technical solutions, review shared-infrastructure changes, and document system behavior and failure modes.

Requirements

  • Bachelor's degree in Computer Science or a similar technical field is required.
  • 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines.
  • Deep production expertise with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
  • Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting; EKS is strongly preferred.
  • Deep experience with RDS/Aurora using MySQL and/or Postgres, including query performance, indexing, connection management, replication, high availability, failover, backup, and recovery.
  • Demonstrated personal delivery of infrastructure scaling improvements with measurable gains in capacity, latency, reliability, or cost efficiency.
  • Strong coding and automation skills using Golang, Python, or similar languages, plus infrastructure-as-code experience with Terraform or an equivalent.
  • Strong systems fundamentals covering Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
  • Experience operating 24/7 high-availability platforms with hands-on incident response and postmortem remediation.
  • Willingness and ability to participate in on-call rotations and recover production systems under pressure.
  • Experience implementing and testing disaster recovery, including data restoration and service-recovery validation.
  • Practical experience with observability, load testing, capacity planning, and safe CI/CD practices for shared infrastructure.
  • Active use of AI tooling in engineering or operations, including verification of generated code, recommendations, and operational actions.
  • Ability to take ambiguous, system-wide technical problems from investigation through production delivery and collaborate across engineering disciplines.
  • Preferred qualifications include fintech, payments, or banking experience; multi-region architectures; chaos engineering; failure testing; distributed data recovery tradeoffs; Prometheus, Grafana, Loki, or Tempo; internal platforms; self-service tooling; progressive delivery; and AI-assisted operational automation.

Benefits

  • Remote, full-time role.
  • Sezzle offers a collaborative culture centered on purpose-driven employees, high standards, innovation, accountability, and measurable results.
  • The company highlights employee interests and activities including music, yoga, cycling, cooking, golf, dogs, and rock climbing.

Sezzle

Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment pr...

Similar positions

Staff Software Engineer - Databases, Tempo

  • Grafana Labs
  • Full time
  • Remote
  • 09/13/2026
  • Salary: €110k-€132k
  • Germany

Software Engineer, Identity Manager

  • Cisco
  • Full time
  • Others
  • 09/13/2026
  • Kraków, Poland

Staff AI Engineer

  • ShiftKey
  • Full time
  • USA
  • 09/13/2026
  • Remote

Principal Software Engineer

  • Array
  • Full time
  • USA
  • 09/13/2026
  • Salary: $200k
  • Remote

Sr. Software Engineer, Backend

  • Scout Motors
  • Full time
  • USA
  • 09/13/2026
  • Salary: $154k - $187k/yr
  • Fremont, CA