Thinking Machines Lab

Thinking Machines Lab

Research Engineer, Infrastructure, Kernels

San Francisco

H1B sponsorship available$350k-$475kDetected 78 days ago
ExpressMachine LearningDeep LearningPyTorchNLPLLMsLLMOpsStatisticsLogisticsRoboticsElectrical EngineeringResearchCommunication

About the role

  • This role is perfect for an engineer who enjoys working close to the metal and across the research boundary.
  • You'll prototype new kernel implementations, profile performance across hardware generations, and help define the numerical and parallelism strategies that determine how we scale next-generation AI systems.
  • Note: This is an "evergreen role" that we keep open on an on-going basis to express interest.

Responsibilities

  • Design and implement custom ML kernels (e.g., CUDA, CuTe, Triton) for core LLM operations such as attention, matrix multiplication, gating, and normalization, optimized for modern GPU and accelerator architectures.
  • Design and think through compute primitives to reduce memory bandwidth bottlenecks and improve kernel compute efficiency.
  • Collaborate with research teams to align kernel-level optimizations with model architecture and algorithmic goals.
  • Develop and maintain a library of reusable kernels and performance benchmarks that serve as the foundation for internal model training.
  • Document and share insights through internal talks, technical papers, or open-source contributions to strengthen the broader ML systems community.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

Requirements

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases
  • Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.

Nice to have

  • we encourage you to apply if you meet some but not all of these:
  • Experience training or supporting large-scale language models with tens of billions of parameters or more.
  • Familiarity with tensor parallelism, pipeline parallelism, or distributed data processing frameworks.
  • Experience implementing low-precision formats (FP8, INT8, block floating point) or contributing to related compiler stacks (e.g., XLA, TVM).
  • Contributions to open-source GPU, ML systems, or compiler optimization projects.
  • Prior research or engineering experience in numerical optimization, communication-efficient training, or scalable AI infrastructure.

Skills

  • Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.

Compensation

  • Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

Benefits

  • Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
  • Bachelor's degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

Visa & Work Authorization

  • While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • We sponsor visas.

This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.