Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Principal Infrastructure Engineer

Home > Golang programming jobs

Principal Infrastructure Engineer in Remote new

  • Brazil

Responsibilities

  • Own the architecture and evolution of core infrastructure as traffic, data volume, and workload complexity grow.
  • Build capacity models, conduct load and stress testing, and diagnose performance bottlenecks across compute, networking, Kubernetes, and databases.
  • Design resilient AWS account, IAM, networking, multi-AZ, and multi-region architectures.
  • Build and operate Kubernetes platforms, including lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
  • Scale and optimize Aurora RDS for MySQL and Postgres through query and index tuning, connection and replication improvements, capacity planning, failover improvements, and safe migrations.
  • Improve reliability with service-level objectives, error budgets, failure isolation, backpressure, load shedding, and safe retry behavior.
  • Participate in on-call rotations and lead technical recovery during major incidents and outages.
  • Design, test, and document disaster recovery, backup, restore, and failover mechanisms against recovery objectives.
  • Build infrastructure-as-code and operational automation for provisioning, configuration, deployment, upgrades, and recovery.
  • Improve observability through metrics, logs, traces, dashboards, and actionable alerts.
  • Deliver safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback paths.
  • Improve cloud cost efficiency through resource sizing, utilization, autoscaling, and storage optimization.
  • Build and evaluate AI-assisted tooling for incident investigation, runbooks, anomaly analysis, and toil reduction with appropriate access controls and auditability.
  • Write architecture proposals, evaluate technology tradeoffs, prototype and benchmark solutions, review infrastructure changes, and document system operations and failure modes.

Requirements

  • Bachelor’s degree in Computer Science or a similar technical field is required.
  • At least 12 years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines is required.
  • Deep production expertise with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
  • Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting; EKS is strongly preferred.
  • Deep experience with RDS/Aurora MySQL and/or Postgres at scale, including query performance, indexing, connection management, replication, high availability, failover, backup, and recovery.
  • Demonstrated personal delivery of infrastructure scaling improvements with measurable gains in capacity, latency, reliability, or cost efficiency.
  • Strong coding and automation skills using Golang, Python, or similar languages, plus infrastructure-as-code experience with Terraform or an equivalent.
  • Strong systems fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
  • Experience operating a 24/7 high-availability platform with hands-on incident response and postmortem remediation.
  • Willingness and ability to participate in on-call rotations and recover production systems under pressure.
  • Experience implementing and testing disaster recovery against defined recovery objectives.
  • Practical experience with observability, load testing, capacity planning, and safe CI/CD practices for shared production infrastructure.
  • Active use of AI tooling in engineering or operations and the judgment to verify generated code, recommendations, and operational actions.
  • Ability to take ambiguous technical problems from investigation through production delivery and collaborate across engineering disciplines.
  • Preferred qualifications include fintech, payments, or banking experience; multi-region architecture, chaos engineering, and failure testing; Prometheus, Grafana, Loki, or Tempo; internal platform and self-service tooling; progressive delivery; and AI-assisted operational automation.

Benefits

  • Full-time remote role.
  • Monthly gross compensation range of $12,500-$20,800, based on location and experience.
  • Opportunity to work on infrastructure supporting fintech and payments workloads with significant reliability, performance, and scale requirements.

Sezzle

Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment pr...

Similar positions

Staff Software Engineer - Databases, Tempo

  • Grafana Labs
  • Full time
  • Remote
  • 09/13/2026
  • Salary: €110k-€132k
  • Germany

Software Engineer, Identity Manager

  • Cisco
  • Full time
  • Others
  • 09/13/2026
  • Kraków, Poland

Staff AI Engineer

  • ShiftKey
  • Full time
  • USA
  • 09/13/2026
  • Remote

Principal Software Engineer

  • Array
  • Full time
  • USA
  • 09/13/2026
  • Salary: $200k
  • Remote

Sr. Software Engineer, Backend

  • Scout Motors
  • Full time
  • USA
  • 09/13/2026
  • Salary: $154k - $187k/yr
  • Fremont, CA