Design and build benchmark suites covering inference performance, model quality, and knowledge evaluation across hardware targets.
Run external partner verifications, compare solutions against benchmarks, identify gaps, and clearly communicate findings.
Port models such as LFM2 across runtimes and frameworks and verify correctness end-to-end.
Maintain and extend the inference engine layer using llama.cpp, ONNX, and MLX as new model architectures emerge.
Make benchmark results explainable, verifiable, reproducible, and independently trustworthy for internal teams and partners.
Requirements
Hands-on experience with at least one inference framework such as llama.cpp, ONNX Runtime, or MLX, including internals and modification beyond basic usage.
Experience designing and building benchmarking pipelines with methodology, validation, and reproducibility.
Strong C++ and Python experience in performance-sensitive contexts.
Solid understanding of inference fundamentals, including quantization, decoding strategies, and memory layout and their interactions.
Preferred experience porting models across runtimes and verifying numerical correctness.
Preferred experience working with external partners or clients in technical validation or evaluation.
Preferred familiarity with edge inference targets and their constraints.
Benefits
Competitive base salary with equity in a unicorn-stage company.
The company pays 100% of medical, dental, and vision premiums for employees and dependents.
401(k) matching of up to 4% of base pay.
Unlimited paid time off and company-wide Refill Days throughout the year.
Full-time, hybrid work arrangement in Boston.
Salary: Competitive
Liquid AI
We build efficient general-purpose AI at every scale.