Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Principal Infrastructure Engineer

Home > Golang programming jobs

Principal Infrastructure Engineer in Remote new

  • France

Responsibilities

  • Own the architecture and evolution of core infrastructure supporting increasing traffic, data volume, and workload complexity.
  • Design resilient AWS account, IAM, networking, service, multi-AZ, and multi-region architectures.
  • Build and operate the Kubernetes platform, including cluster lifecycle, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
  • Scale and optimize Aurora RDS for MySQL and Postgres through query and index tuning, connection and replication improvements, capacity planning, failover, schema changes, and migrations.
  • Engineer reliability through service-level objectives, error budgets, failure isolation, backpressure, load shedding, and safe retry behavior.
  • Participate in on-call rotations and lead technical recovery during serious incidents, including full outages.
  • Design, test, and document disaster recovery, backup, restore, failover, and recovery exercises.
  • Build infrastructure-as-code, deployment, provisioning, configuration, upgrade, and operational automation.
  • Improve observability through metrics, logs, traces, dashboards, and actionable alerts.
  • Deliver safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback paths.
  • Improve cloud cost efficiency through resource right-sizing, utilization improvements, autoscaling, storage tuning, and quantified savings.
  • Build and evaluate AI-assisted tooling for incident investigation, capacity analysis, runbook automation, anomaly analysis, and toil reduction.
  • Write architecture proposals, evaluate tradeoffs through prototypes and benchmarks, review shared-infrastructure changes, and document system operation and failure modes.

Requirements

  • Bachelor’s degree in Computer Science or a similar technical field is required.
  • 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines is required.
  • Deep production expertise with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
  • Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting; EKS experience is preferred.
  • Deep expertise with RDS/Aurora MySQL and/or Postgres, including query performance, indexing, connection management, replication, high availability, failover, backup, and recovery.
  • Demonstrated personal delivery of infrastructure scaling improvements with measurable capacity, latency, reliability, or cost gains.
  • Strong coding and automation skills with Golang, Python, or similar languages, plus infrastructure-as-code experience with Terraform or equivalent.
  • Strong systems fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
  • Experience operating a 24/7 high-availability platform with hands-on incident response and postmortem remediation.
  • Willingness and demonstrated ability to participate in on-call rotations and recover production systems under pressure.
  • Experience implementing and testing disaster recovery against defined recovery objectives.
  • Practical experience with observability, load testing, capacity planning, and safe CI/CD practices for shared production infrastructure.
  • Active use of AI tooling in engineering or operations, with the ability to verify generated code, recommendations, and operational actions.
  • Ability to take ambiguous technical problems from investigation through production delivery and collaborate across engineering disciplines.
  • Preferred qualifications include fintech, payments, or banking experience; multi-region architectures; chaos engineering; failure testing; Prometheus, Grafana, Loki, or Tempo; internal platform and self-service tooling; progressive delivery; and AI-assisted incident investigation or operational automation.

Benefits

  • Remote, full-time position.
  • Monthly gross compensation range of $12,500–$20,800 USD based on location and experience.
  • Opportunity to work on infrastructure supporting fintech and payments workloads with direct customer and revenue impact.
  • Collaborative culture with engineers, data enthusiasts, innovators, and colleagues with interests including music, yoga, cycling, cooking, golf, dogs, and rock climbing.

Sezzle

Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment pr...

Similar positions

Senior Software Engineer - Integrations 3P

  • Elastic
  • Full time
  • Remote
  • 09/16/2026
  • Salary: €61k-€97k
  • Greece

Staff Software Engineer, Backend

  • Rivian and Volkswagen Group Technologies
  • Full time
  • USA
  • 09/16/2026
  • Salary: $186k - $255k/yr
  • Palo Alto, CA

Senior Software Engineer - Integrations 3P

  • Elastic
  • Full time
  • Remote
  • 09/16/2026
  • Salary: €87k-€138k
  • Ireland

Senior Software Engineer - Integrations 3P

  • Elastic
  • Full time
  • Remote
  • 09/16/2026
  • Salary: €67k-€106k
  • Spain

Staff Software Engineer - Databases, Tempo

  • Grafana Labs
  • Full time
  • UK
  • 09/15/2026
  • Salary: £104k-£125k
  • Remote