Thinking Machines Lab
Research Engineer, Infrastructure, Inference
San Francisco
H1B sponsorship available$350k-$475kDetected 5 days ago
ExpressKubernetesMachine LearningDeep LearningPyTorchAI OrchestrationLogisticsResearch
About the role
- Your work will make inference faster, more cost-effective, more reliable, and more reproducible to enable our teams to focus on advancing model capabilities rather than managing bottlenecks.
- Our focus is on performant and efficient model inference both to power real-world applications and to accelerate research.
- This role is responsible for the infrastructure that ensures every experiment, evaluation, and deployment runs smoothly at scale.
Responsibilities
- Collaborate with research teams to enable high-performance inference for novel architectures.
- Design and implement new techniques, tools, and architectures that improve performance, latency, throughput, and efficiency.
- Optimize our codebase and compute fleet (e.g., GPUs) to fully utilize hardware FLOPs, bandwidth, and memory.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, or similar.
- Experience with inference serving systems optimized for throughput and latency (e.g., SGLang, vLLM).
- Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases
Nice to have
- we encourage you to apply if you meet some but not all of these:
- Experience training or supporting large-scale language models with hundreds of billions of parameters or more.
- Understanding of distributed compute systems, GPU parallelism, and hardware-aware optimizations.
- Contributions to open-source ML or systems infrastructure projects (e.g., SGLang, vLLM, PyTorch, Triton, DeepSpeed, XLA).
Skills
- Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
Compensation
- Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
Benefits
- Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
Visa & Work Authorization
- While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- We sponsor visas.
Apply directly at Thinking Machines Lab →Create a free account for alerts like thisView Thinking Machines Lab immigration profile
This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.