Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Principal Infrastructure Engineer

Home > Golang programming jobs

Principal Infrastructure Engineer in Remote new

  • Turkey

Responsibilities

  • Own the technical architecture and evolution of core infrastructure as traffic, data volume, and workload complexity increase.
  • Build capacity models, conduct load and stress testing, diagnose bottlenecks, and improve throughput, latency, reliability, and cost efficiency.
  • Design resilient AWS account, IAM, networking, multi-AZ, and multi-region architectures.
  • Build and operate the Kubernetes platform, including cluster lifecycle, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
  • Scale and optimize Aurora RDS for MySQL and Postgres through query and index tuning, connection and replication improvements, capacity planning, failover improvements, and safe migrations.
  • Define and implement reliability practices including service-level objectives, error budgets, failure isolation, backpressure, load shedding, and safe retries.
  • Participate in on-call rotations, respond to major incidents, execute mitigations, and implement postmortem corrective actions.
  • Design, test, and document disaster recovery, backup, restore, and failover mechanisms against recovery objectives.
  • Build infrastructure-as-code, deployment, provisioning, upgrade, recovery, and operational automation.
  • Improve observability through metrics, logs, traces, dashboards, and actionable alerts.
  • Deliver safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback paths.
  • Improve cloud cost efficiency through resource right-sizing, utilization improvements, autoscaling, and storage optimization.
  • Build and evaluate AI-assisted tooling for incident investigation, runbooks, anomaly analysis, and repetitive operations.
  • Write architecture proposals, evaluate technical tradeoffs with prototypes and benchmarks, and document system behavior and failure modes.

Requirements

  • Bachelor’s degree in Computer Science or a similar technical field is required.
  • 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines is required.
  • Deep production expertise with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
  • Deep production expertise with Kubernetes; EKS experience is strongly preferred.
  • Deep expertise with RDS/Aurora using MySQL and/or Postgres, including performance, indexing, connections, replication, high availability, failover, backup, and recovery.
  • Demonstrated delivery of infrastructure scaling improvements with measurable capacity, latency, reliability, or cost results.
  • Strong coding and automation skills using Golang, Python, or similar languages, plus Terraform or equivalent infrastructure-as-code experience.
  • Strong fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
  • Experience operating 24/7 high-availability platforms with hands-on incident response and postmortem remediation.
  • Willingness and ability to participate in on-call rotations and recover production systems under pressure.
  • Experience implementing and testing disaster recovery against defined recovery objectives.
  • Practical experience with observability, load testing, capacity planning, and safe CI/CD practices.
  • Active use of AI tooling in engineering or operations, including verification of generated code, recommendations, and operational actions.
  • Ability to take ambiguous technical problems from investigation through production delivery and collaborate across engineering disciplines.
  • Preferred experience includes fintech, payments, or banking; multi-region architectures; chaos engineering; failure testing; Prometheus, Grafana, Loki, or Tempo; internal platforms; self-service tooling; progressive delivery; and AI-assisted incident automation.

Benefits

  • Full-time remote role, indicated by the #Li-remote and #full-time tags.
  • Monthly gross compensation range of $12,500–$20,800 based on location and experience.
  • Collaborative culture with engineers, data enthusiasts, innovators, and colleagues with varied personal interests and activities.

Sezzle

Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment pr...

Similar positions

Staff Software Engineer - Databases, Tempo

  • Grafana Labs
  • Full time
  • Remote
  • 09/13/2026
  • Salary: €110k-€132k
  • Germany

Software Engineer, Identity Manager

  • Cisco
  • Full time
  • Others
  • 09/13/2026
  • Kraków, Poland

Staff AI Engineer

  • ShiftKey
  • Full time
  • USA
  • 09/13/2026
  • Remote

Principal Software Engineer

  • Array
  • Full time
  • USA
  • 09/13/2026
  • Salary: $200k
  • Remote

Sr. Software Engineer, Backend

  • Scout Motors
  • Full time
  • USA
  • 09/13/2026
  • Salary: $154k - $187k/yr
  • Fremont, CA