Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Senior AI Training Infrastructure Engineer

Home > Javascript/Typescript programming jobs

Senior AI Training Infrastructure Engineer in USA new

  • Designworks Talent
  • Full time
  • Email
  • Bellevue, WA

Responsibilities

  • Build and scale distributed training infrastructure for large AI models across GPU clusters.
  • Design systems that improve training reliability, efficiency, fault tolerance, checkpointing, recovery, and resource utilization.
  • Integrate AI models into production training pipelines with platform, orchestration, and performance engineering teams.
  • Diagnose issues affecting training throughput, stability, reliability, and cost efficiency.
  • Build automation and developer tools for AI researchers and engineers.
  • Establish best practices for training infrastructure, operational processes, and platform reliability.
  • Contribute to the architecture and evolution of the AI infrastructure platform.

Requirements

  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.
  • Strong understanding of reliability, scalability, and efficiency challenges in multi-node GPU training.
  • Experience integrating training systems with production machine learning pipelines.
  • Strong programming skills and experience with complex distributed systems.
  • Ability to independently own technically challenging projects in a fast-moving, high-ownership environment.
  • Preferred experience with PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar distributed training technologies.
  • Preferred experience with supervised fine-tuning, reinforcement learning from human feedback, GPU optimization, training performance, or distributed-system reliability.
  • Preferred background operating AI training infrastructure at scale in a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.
  • Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms.

Benefits

  • Hybrid role in downtown Bellevue, Washington, with approximately three days per week in the office.
  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.
  • Medical, dental, and vision insurance for U.S.-based employees.
  • 401(k) plan with company match.
  • Paid holidays per calendar year.
  • Certain roles may be eligible for merit increases, annual bonuses, and long-term incentives.

Designworks Talent

Designworks Talent specializes in the design and delivery of enterprise workforce solutions including strategy, talent acquisition, engagement, and succession planning. We work collaboratively with business leaders and enterprise Talent Acquisition and Human Resources teams to deliver custom work...

Similar positions

Frontend Engineer - Mobile

  • Betr
  • Full time
  • Canada
  • 10/04/2026
  • Remote

Senior Software Engineer, Frontend

  • Circle
  • Full time
  • USA
  • 10/04/2026
  • Salary: $152k-$205k
  • Remote

Senior Platform Engineer

  • Clera
  • Full time
  • Others
  • 10/04/2026
  • Munich, Germany

Full-Stack NodeJS Software Engineer

  • PeopleGrove
  • Full time
  • Remote
  • 10/04/2026
  • Salary: Competitive
  • India

Staff Software Engineer

  • G-P
  • Full time
  • Remote
  • 10/03/2026
  • Salary: £70k-£97k
  • UK, Ireland