Thinking Machines Lab
Research Engineer, Infrastructure, Numerics
San Francisco
H1B sponsorship available$350k-$475kDetected 78 days ago
Node.jsExpressMachine LearningDeep LearningPyTorchLLMsStatisticsLogisticsSystems EngineeringRoboticsElectrical EngineeringResearchCommunication
About the role
- You will focus on improving the numerical foundations of our distributed training stack, from precision formats and kernel optimizations to communication frameworks that make training trillion-parameter models stable, scalable, and fast.
- This role is ideal for someone who thrives at the intersection of research and systems engineering: a builder who understands both the math of optimization and the realities of distributed compute.
- Note: This is an "evergreen role" that we keep open on an on-going basis to express interest.
Responsibilities
- Design and optimize distributed training infrastructure for large-scale LLMs, focusing on performance, stability, and reproducibility across multi-GPU and multi-node setups.
- Implement and evaluate low-precision numerics (for example, BF16, MXFP8, NVFP4) to improve efficiency without sacrificing model quality.
- Develop kernels and communication primitives that use hardware-level support for mixed and low-precision arithmetic.
- Collaborate with research teams to co-design model architectures and training recipes that align with emerging numeric formats and stability constraints.
- Contribute to the design of our internal orchestration and monitoring systems to ensure that thousands of distributed experiments can run efficiently and reproducibly.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
Requirements
- Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases in areas such as floating-point numerics, low-precision arithmetic, and distributed systems.
Nice to have
- we encourage you to apply if you meet some but not all of these:
- Familiarity with distributed frameworks such as PyTorch/XLA, DeepSpeed, Megatron-LM.
- Experience implementing FP8, INT8, or block-floating point (MX) formats and understanding their numerical trade-offs.
- Publications, patents, or projects related to numerical optimization, communication-efficient training, or systems for large models.
- Experience training and supporting large-scale AI models.
Skills
- Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
Compensation
- Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
Benefits
- Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
- Bachelor's degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
Visa & Work Authorization
- While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- We sponsor visas.
Apply directly at Thinking Machines Lab →Create a free account for alerts like thisView Thinking Machines Lab immigration profile
This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.