Inference
Senior Software Engineer - Model Performance
San Francisco · Senior
Sponsorship not specified$220k-$320kDetected 181 days ago
PythonC++DockerKubernetesMachine LearningPyTorchNLPLLMsCustomer SupportAdaptability
About the role
- You will be responsible for making our inference stack as fast and efficient as possible.
- Your work spans from implementing known optimization techniques to experimenting with novel approaches, always with the goal of serving models faster and cheaper at scale.
- Your north star is inference performance: latency, throughput, cost efficiency, and how quickly we can bring new model architectures into production.
Responsibilities
- Implement and productionize optimization techniques including quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving
- Profile and optimize CUDA kernels and GPU utilization across our serving infrastructure
- Add support for new model architectures, ensuring they meet our performance standards before going to production
- Build tooling and benchmarks to measure and track inference performance across our fleet
- Collaborate with applied ML engineers to ensure trained models can be served efficiently
Requirements
- 2+ years of experience in ML systems, inference optimization, or GPU programming
- Strong proficiency in Python and familiarity with C++
- Hands-on experience with LLM inference frameworks (vLLM, SGLang, TensorRT-LLM, or similar)
- Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching, KV cache management)
- Experience with PyTorch and understanding of how models execute on hardware
- Track record of measurably improving system performance
Nice to have
- Experience with CUDA programming
- Experience with distributed inference and multi-GPU serving
- Contributions to open-source inference frameworks
- Experience with Docker and Kubernetes
- You don't need to tick every box.
- Curiosity and the ability to learn quickly matter more.
Skills
- Deep dive into inference frameworks (vLLM, SGLang, TensorRT-LLM) and underlying libraries to debug and improve performance
- Experiment with novel inference techniques and bring successful approaches into production
Compensation
- We offer competitive compensation, equity in a high-growth startup, and comprehensive benefits.
- The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience.
Equal opportunity
- Inference.net http://Inference.net is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.
- If you're excited about making AI inference faster for everyone, we'd love to hear from you. Please send your resume and GitHub to amar@inference.net and/or apply here on Ashby.
- Inference.net http://Inference.net is an equal opportunity employer.
- We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.
Apply directly at Inference →Create a free account for alerts like thisView Inference immigration profile
This listing is sourced directly from Inference's careers page and normalized into a canonical job model.