Train and optimize large-scale video and multimodal models
Improve efficiency across training and inference, including memory, latency, and cost
Implement distillation, quantization, and pruning techniques to accelerate diffusion and autoregressive generation
Build and maintain distributed training systems
Optimize GPU utilization, parallelism, and throughput
Develop tooling for experimentation, evaluation, and debugging
Translate research models into robust, production-ready systems
Monitor and improve model performance in real-world usage
Requirements
BS, MS, or PhD in computer science, machine learning, or a related field
At least 2 years of professional industry experience
Strong experience with deep learning systems and infrastructure
Expertise in PyTorch, CUDA, Triton, and distributed training, including FSDP
Experience scaling and optimizing large models under low-latency inference constraints
Strong debugging and performance profiling skills
Ability to move quickly from prototype to production
Benefits
Comprehensive medical, dental, and vision plans
401K with employer match
Commuter benefits
Catered lunch multiple days per week
Dinner stipend every night when working late
Grubhub subscription
Health and wellness perks
Multiple team offsites per year and monthly team events
Generous PTO policy
In-person work required at the NYC headquarters in Union Square
Salary: $175k - $275k/yr
Mirage
At Mirage, we’re building full-stack foundation models and products that redefine video creation. Over 20 million content creators and businesses use Captions to reach their full creative and commercial potential.