Thinking Machines Lab
Research Engineer, Infrastructure, Training Systems
San Francisco
H1B sponsorship available$350k-$475kDetected 78 days ago
ExpressMachine LearningDeep LearningPyTorchStatisticsA/B TestingLogisticsRoboticsElectrical EngineeringResearch
About the role
- Your goal is to make experimentation and training at Thinking Machines fast and reliable to ensure our research teams can focus on science, not system bottlenecks.
- This role is ideal for someone who blends deep systems and performance expertise with a curiosity for machine learning at scale.
- Note: This is an "evergreen role" that we keep open on an on-going basis to express interest.
Responsibilities
- Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes for large-scale training workloads.
- Develop high-performance optimizations to maximize throughput and efficiency.
- Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.
- Collaborate with researchers and engineers to build scalable infrastructure.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
Requirements
- Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases
Nice to have
- we encourage you to apply if you meet some but not all of these:
- Past experience working on distributed training for the world's largest models to make them stable, reliable, and performant.
- Contributions to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.
Skills
- Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
Compensation
- Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
Benefits
- Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
- Bachelor's degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
Visa & Work Authorization
- While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- We sponsor visas.
Apply directly at Thinking Machines Lab →Create a free account for alerts like thisView Thinking Machines Lab immigration profile
This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.