Thinking Machines Lab
Research Engineer, Infrastructure, Kernels
San Francisco
H1B sponsorship available$350k-$475kDetected 78 days ago
ExpressMachine LearningDeep LearningPyTorchNLPLLMsLLMOpsStatisticsLogisticsRoboticsElectrical EngineeringResearchCommunication
About the role
- This role is perfect for an engineer who enjoys working close to the metal and across the research boundary.
- You'll prototype new kernel implementations, profile performance across hardware generations, and help define the numerical and parallelism strategies that determine how we scale next-generation AI systems.
- Note: This is an "evergreen role" that we keep open on an on-going basis to express interest.
Responsibilities
- Design and implement custom ML kernels (e.g., CUDA, CuTe, Triton) for core LLM operations such as attention, matrix multiplication, gating, and normalization, optimized for modern GPU and accelerator architectures.
- Design and think through compute primitives to reduce memory bandwidth bottlenecks and improve kernel compute efficiency.
- Collaborate with research teams to align kernel-level optimizations with model architecture and algorithmic goals.
- Develop and maintain a library of reusable kernels and performance benchmarks that serve as the foundation for internal model training.
- Document and share insights through internal talks, technical papers, or open-source contributions to strengthen the broader ML systems community.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
Requirements
- Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases
- Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
Nice to have
- we encourage you to apply if you meet some but not all of these:
- Experience training or supporting large-scale language models with tens of billions of parameters or more.
- Familiarity with tensor parallelism, pipeline parallelism, or distributed data processing frameworks.
- Experience implementing low-precision formats (FP8, INT8, block floating point) or contributing to related compiler stacks (e.g., XLA, TVM).
- Contributions to open-source GPU, ML systems, or compiler optimization projects.
- Prior research or engineering experience in numerical optimization, communication-efficient training, or scalable AI infrastructure.
Skills
- Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
Compensation
- Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
Benefits
- Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
- Bachelor's degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
Visa & Work Authorization
- While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- We sponsor visas.
Apply directly at Thinking Machines Lab →Create a free account for alerts like thisView Thinking Machines Lab immigration profile
This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.