Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Principal Infrastructure Engineer

Home > Golang programming jobs

Principal Infrastructure Engineer in Remote new

  • Argentina

Responsibilities

  • Own the technical architecture and evolution of core infrastructure as traffic, data volume, and workload complexity grow.
  • Build capacity models, conduct load and stress tests, diagnose bottlenecks, and improve throughput, latency, reliability, and cost efficiency.
  • Design resilient AWS account, IAM, networking, multi-AZ, and multi-region architectures.
  • Build and operate Kubernetes platforms, including cluster lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
  • Scale and optimize Aurora RDS databases using MySQL and Postgres, including query tuning, indexing, connection management, replication, failover, schema changes, and migrations.
  • Define and instrument service-level objectives and error budgets and implement failure isolation, backpressure, load shedding, and safe retry behavior.
  • Participate in on-call rotations, recover major incidents, perform evidence-based triage, and implement postmortem corrective actions.
  • Design, test, and document disaster recovery, backup, restore, and failover mechanisms against recovery objectives.
  • Build infrastructure-as-code, provisioning, deployment, upgrade, recovery, observability, and operational automation.
  • Deliver safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback paths.
  • Improve cloud cost efficiency through resource sizing, utilization improvements, autoscaling, and storage optimization.
  • Build and evaluate AI-assisted tooling for incident investigation, runbooks, anomaly analysis, capacity analysis, and toil reduction.
  • Write architecture proposals, prototype and benchmark technology choices, review shared infrastructure changes, and document system operations and failure modes.

Requirements

  • Bachelor’s degree in Computer Science or a similar technical field is required.
  • 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines.
  • Deep production expertise with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
  • Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting; EKS is strongly preferred.
  • Deep expertise with relational databases at scale, specifically RDS/Aurora with MySQL and/or Postgres.
  • Demonstrated experience personally delivering infrastructure scaling improvements with measurable gains in capacity, latency, reliability, or cost efficiency.
  • Strong coding and automation skills using Golang, Python, or similar languages, plus infrastructure-as-code experience with Terraform or equivalent.
  • Strong systems fundamentals covering Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
  • Experience operating 24/7 high-availability platforms with hands-on incident response and postmortem remediation.
  • Willingness and ability to participate in on-call rotations and recover production systems under pressure.
  • Experience implementing and testing disaster recovery, restoring data, and validating service recovery.
  • Practical experience with observability, load testing, capacity planning, and safe CI/CD practices.
  • Active use of AI tooling in engineering or operations, including verification of generated code, recommendations, and operational actions.
  • Ability to carry ambiguous technical problems through investigation, production delivery, and cross-disciplinary resolution.
  • Preferred experience includes fintech, payments, or banking; multi-region architectures; chaos engineering; failure testing; Prometheus, Grafana, Loki, or Tempo; and internal platform and self-service tooling.

Benefits

  • Remote, full-time position.
  • Employees participate in a culture that includes musicians, yogis, cyclists, chefs, golfers, dog-lovers, and rock-climbers.

Sezzle

Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment pr...

Similar positions

Staff Software Engineer - Databases, Tempo

  • Grafana Labs
  • Full time
  • Remote
  • 09/13/2026
  • Salary: €110k-€132k
  • Germany

Software Engineer, Identity Manager

  • Cisco
  • Full time
  • Others
  • 09/13/2026
  • Kraków, Poland

Staff AI Engineer

  • ShiftKey
  • Full time
  • USA
  • 09/13/2026
  • Remote

Principal Software Engineer

  • Array
  • Full time
  • USA
  • 09/13/2026
  • Salary: $200k
  • Remote

Sr. Software Engineer, Backend

  • Scout Motors
  • Full time
  • USA
  • 09/13/2026
  • Salary: $154k - $187k/yr
  • Fremont, CA