Thinking Machines Lab

Thinking Machines Lab

Research Engineer, Infrastructure, Inference

San Francisco

H1B sponsorship available$350k-$475kDetected 5 days ago
ExpressKubernetesMachine LearningDeep LearningPyTorchAI OrchestrationLogisticsResearch

About the role

  • Your work will make inference faster, more cost-effective, more reliable, and more reproducible to enable our teams to focus on advancing model capabilities rather than managing bottlenecks.
  • Our focus is on performant and efficient model inference both to power real-world applications and to accelerate research.
  • This role is responsible for the infrastructure that ensures every experiment, evaluation, and deployment runs smoothly at scale.

Responsibilities

  • Collaborate with research teams to enable high-performance inference for novel architectures.
  • Design and implement new techniques, tools, and architectures that improve performance, latency, throughput, and efficiency.
  • Optimize our codebase and compute fleet (e.g., GPUs) to fully utilize hardware FLOPs, bandwidth, and memory.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Experience with inference serving systems optimized for throughput and latency (e.g., SGLang, vLLM).
  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases

Nice to have

  • we encourage you to apply if you meet some but not all of these:
  • Experience training or supporting large-scale language models with hundreds of billions of parameters or more.
  • Understanding of distributed compute systems, GPU parallelism, and hardware-aware optimizations.
  • Contributions to open-source ML or systems infrastructure projects (e.g., SGLang, vLLM, PyTorch, Triton, DeepSpeed, XLA).

Skills

  • Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.

Compensation

  • Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

Benefits

  • Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

Visa & Work Authorization

  • While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • We sponsor visas.

This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.