Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Sr. Software Engineer, AI Infrastructure

Home > Python programming jobs

Sr. Software Engineer, AI Infrastructure in USA new

  • Washington, DC

Responsibilities

  • Manage GPU and CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU-as-a-service support for external customers on bare-metal and virtualized platforms.
  • Design, validate, and productize AI cluster solutions at 100,000-plus GPU scale.
  • Develop automation for on-premise Kubernetes and AI clusters and operating systems.
  • Deploy and manage databases, monitoring systems, and distributed storage.
  • Collaborate with AI engineers to build scalable, operable, and maintainable products.
  • Manage the full service lifecycle from design and deployment through operation and refinement.
  • Improve monitoring, alerting, and system availability.
  • Mentor and train junior engineers and lead the team toward technical excellence.
  • Work extended hours and weekends and travel domestically and globally as needed.

Requirements

  • Bachelor’s degree in computer science, information systems/IT, or engineering plus 5 or more years of professional Linux experience, or 7 or more years of professional software, DevOps, or site reliability engineering experience in lieu of a degree.
  • At least 5 years of Kubernetes experience and at least 5 years managing Linux operating systems.
  • Experience with Terraform, Ansible, or comparable infrastructure tools and with containerization technologies such as OCI containers and Kubernetes.
  • Scripting experience in Bash, Python, or similar languages, plus development experience in Python, C++, or Go.
  • Preferred experience includes Python-based development frameworks, Kubernetes cluster management, Linux boot and systems configuration, testing, continuous integration, build and deployment technologies, and continuous monitoring.
  • Preferred qualifications include Bazel or Makefiles, performance optimization, distributed databases and data modeling, large-scale server automation, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks including Blackwell and Rubin.
  • Active Top Secret, Top Secret SCI, or DOE Level Q clearance is preferred.
  • Must successfully obtain and maintain a Top Secret Security Clearance as a condition of employment.
  • Must meet ITAR eligibility requirements or be eligible to obtain the required U.S. Department of State authorizations.

Benefits

  • Base salary range is $165,000.00-$265,000.00 annually for Level 3.
  • Eligible employees may receive long-term incentives, potential discretionary bonuses, and Employee Stock Purchase Plan access.
  • Benefits include medical, vision, dental, 401(k), disability and life insurance, paid parental leave, discounts, paid vacation, paid holidays, and paid sick leave.
  • Employees with an active clearance may receive a 10% differential up to an additional $20,000 annually after being briefed into a classified program.
  • The role may require extended hours, weekend work, and domestic or global travel.

Salary: $165k - $265k/yr

SpaceX

SpaceX builds advanced rockets and spacecraft with the aim of making human life multiplanetary, particularly focusing on establishing a sustainable presence on Mars. Its innovative technologies and ambitious missions cater to space enthusiasts, researchers, and industries looking to push the boun...

Similar positions

GTM AI Automation Engineer

  • Factored
  • Full time
  • Remote
  • 10/02/2026
  • Americas

Senior Staff Engineer, Product Security

  • Druva
  • Full time
  • Others
  • 10/02/2026
  • Pune, India

Applied AI Product Engineer

  • VetsEZ
  • Full time
  • USA
  • 10/02/2026
  • Remote

AI Security Researcher/Research Engineer

  • Brave
  • Full time
  • Remote
  • 10/02/2026
  • Salary: Competitive
  • Europe, North America

Python and Kubernetes Software Engineer

  • Canonical
  • Full time
  • Remote
  • 10/02/2026
  • Anywhere