Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Inference Systems Performance Architect

Home > Other programming jobs

Inference Systems Performance Architect in USA

  • SambaNova Systems
  • Full time
  • Email
  • San Jose, CA

Responsibilities

  • Define and drive the technical strategy for inference-systems performance across workload capture, benchmarking, modeling, and simulation.
  • Build workload-capture and agentic-benchmarking capabilities that represent production traffic and identify artificial contention or misleading cache-hit rates.
  • Own performance modeling and simulation practices to inform capacity planning, customer SLOs, and future systems and hardware planning.
  • Develop profiling tools that accurately localize bottlenecks across distributed inference hosts, accelerators, and fabrics.
  • Act as the senior technical voice across model optimization, systems, hardware, and product teams while evaluating reliability, scalability, operational cost, and adoption trade-offs.
  • Represent SambaNova’s performance capabilities to customers and partners.
  • Mentor principal and senior engineers and create systems, tools, and patterns that improve organizational productivity.
  • Drive resolution of ambiguous, novel, cross-organizational performance challenges.

Requirements

  • 12+ years of experience in performance engineering with a record of technical leadership on large-scale, complex systems.
  • Deep expertise in end-to-end performance analysis of distributed systems and bottleneck localization.
  • Proven expertise in realistic workload generation, simulation, and performance modeling calibrated against variable real-world workloads.
  • Ability to apply transferable performance methods in unfamiliar domains.
  • Ability to lead cross-functional efforts, mentor senior engineers, and influence organizational direction.
  • Experience credibly representing an organization to customers and partners.
  • Track record of independently scoping and delivering high-complexity, high-ambiguity work with significant product or roadmap impact.
  • Preferred: direct experience with LLM inference serving, continuous batching, prompt/KV caching, prefill/decode disaggregation, and tail-latency SLOs.
  • Preferred: familiarity with inference simulation frameworks or agentic benchmarking efforts.
  • Preferred: public technical presence through talks, writing, or community participation on systems performance.

Benefits

  • Base benefits include medical insurance with 95% employee premium coverage and 77% dependent premium coverage, plus an employer-contributed Health Savings Account.
  • Dental, vision, short- and long-term disability, basic life, voluntary life, AD&D, and flexible spending account options are available.
  • Well-being benefits include Headspace, Gympass+ with access to physical gyms, One Medical, counseling services, and an Employee Assistance Program.
  • The role is a full-time position for US-based employment.

Salary: $245k - $325k/yr

SambaNova Systems

Welcome to SambaNova: Revolutionizing AI Capacity At SambaNova, we're empowering developers, enterprises, governments, and data centers to unlock their full AI potential. Our full-stack infrastructure, from chips to models, enables lightning-fast performance, low power consumption, and high-effic...

Similar positions

INFRASTRUCTURE AND PLATFORM ARCHITECT L2

  • Wipro
  • Full time
  • Others
  • 09/13/2026
  • Bengaluru, India

Salesforce Architect

  • Rula
  • Full time
  • USA
  • 09/13/2026
  • Salary: $152k - $177k/yr
  • Remote

AI Engineer, Internal Systems

  • Wispr Flow
  • Full time
  • USA
  • 09/12/2026
  • Salary: $220k - $300k/yr
  • Remote

Enterprise Architect

  • Blackbaud
  • Full time
  • USA
  • 09/12/2026
  • Salary: Competitive
  • Remote

Solution Architect - AI

  • Litify
  • Full time
  • USA
  • 09/12/2026
  • Salary: $120k
  • Remote