Thinking Machines Lab

Thinking Machines Lab

Research Engineer, Infrastructure, Numerics

San Francisco

H1B sponsorship available$350k-$475kDetected 78 days ago
Node.jsExpressMachine LearningDeep LearningPyTorchLLMsStatisticsLogisticsSystems EngineeringRoboticsElectrical EngineeringResearchCommunication

About the role

  • You will focus on improving the numerical foundations of our distributed training stack, from precision formats and kernel optimizations to communication frameworks that make training trillion-parameter models stable, scalable, and fast.
  • This role is ideal for someone who thrives at the intersection of research and systems engineering: a builder who understands both the math of optimization and the realities of distributed compute.
  • Note: This is an "evergreen role" that we keep open on an on-going basis to express interest.

Responsibilities

  • Design and optimize distributed training infrastructure for large-scale LLMs, focusing on performance, stability, and reproducibility across multi-GPU and multi-node setups.
  • Implement and evaluate low-precision numerics (for example, BF16, MXFP8, NVFP4) to improve efficiency without sacrificing model quality.
  • Develop kernels and communication primitives that use hardware-level support for mixed and low-precision arithmetic.
  • Collaborate with research teams to co-design model architectures and training recipes that align with emerging numeric formats and stability constraints.
  • Contribute to the design of our internal orchestration and monitoring systems to ensure that thousands of distributed experiments can run efficiently and reproducibly.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

Requirements

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases in areas such as floating-point numerics, low-precision arithmetic, and distributed systems.

Nice to have

  • we encourage you to apply if you meet some but not all of these:
  • Familiarity with distributed frameworks such as PyTorch/XLA, DeepSpeed, Megatron-LM.
  • Experience implementing FP8, INT8, or block-floating point (MX) formats and understanding their numerical trade-offs.
  • Publications, patents, or projects related to numerical optimization, communication-efficient training, or systems for large models.
  • Experience training and supporting large-scale AI models.

Skills

  • Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.

Compensation

  • Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

Benefits

  • Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
  • Bachelor's degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

Visa & Work Authorization

  • While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • We sponsor visas.

This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.