Thinking Machines Lab

Thinking Machines Lab

Research Engineer, Infrastructure, RL Systems

San Francisco

H1B sponsorship available$350k-$475kDetected 78 days ago
ExpressAlgorithmsKubernetesPrometheusGrafanaMachine LearningDeep LearningPyTorchStatisticsLogisticsSystems EngineeringRoboticsElectrical EngineeringResearch

About the role

  • This role sits at the intersection of research and large-scale systems engineering: a builder who understands both the algorithms behind RL and the realities of distributed training and inference at scale.
  • You'll wear many hats, from optimizing rollout and reward pipelines to enhancing reliability, observability, and orchestration, collaborating closely with researchers and infra teams to make reinforcement learning stable, fast, and production-ready.
  • Note: This is an "evergreen role" that we keep open on an on-going basis to express interest.

Responsibilities

  • Develop shared monitoring and observability tools to ensure high uptime, debuggability, and reproducibility for RL systems.
  • Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines.
  • Build evaluation and benchmarking infrastructure that measures model progress on helpfulness, safety, and factuality.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

Requirements

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases

Nice to have

  • we encourage you to apply if you meet some but not all of these:
  • Experience training or supporting large-scale language models with tens of billions of parameters or more.
  • Background in high-performance or reliability engineering - distributed training frameworks and cluster orchestration (Kubernetes, Slurm).
  • Familiarity with monitoring and observability tools (Prometheus, Grafana, OpenTelemetry).
  • Contributions to large-scale ML research or infrastructure, open-source frameworks, or internal performance optimization efforts.

Skills

  • Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.

Compensation

  • Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

Benefits

  • Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
  • Design, build, and optimize the infrastructure that powers large-scale reinforcement learning and post-training workloads.
  • Bachelor's degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

Visa & Work Authorization

  • While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • We sponsor visas.

This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.