Thinking Machines Lab

Thinking Machines Lab

Research Engineer, Infrastructure, Training Systems

San Francisco

H1B sponsorship available$350k-$475kDetected 78 days ago
ExpressMachine LearningDeep LearningPyTorchStatisticsA/B TestingLogisticsRoboticsElectrical EngineeringResearch

About the role

  • Your goal is to make experimentation and training at Thinking Machines fast and reliable to ensure our research teams can focus on science, not system bottlenecks.
  • This role is ideal for someone who blends deep systems and performance expertise with a curiosity for machine learning at scale.
  • Note: This is an "evergreen role" that we keep open on an on-going basis to express interest.

Responsibilities

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes for large-scale training workloads.
  • Develop high-performance optimizations to maximize throughput and efficiency.
  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.
  • Collaborate with researchers and engineers to build scalable infrastructure.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

Requirements

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases

Nice to have

  • we encourage you to apply if you meet some but not all of these:
  • Past experience working on distributed training for the world's largest models to make them stable, reliable, and performant.
  • Contributions to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.

Skills

  • Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.

Compensation

  • Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

Benefits

  • Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
  • Bachelor's degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

Visa & Work Authorization

  • While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • We sponsor visas.

This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.