Thinking Machines Lab
Research Engineer, Infrastructure, RL Systems
San Francisco
H1B sponsorship available$350k-$475kDetected 78 days ago
ExpressAlgorithmsKubernetesPrometheusGrafanaMachine LearningDeep LearningPyTorchStatisticsLogisticsSystems EngineeringRoboticsElectrical EngineeringResearch
About the role
- This role sits at the intersection of research and large-scale systems engineering: a builder who understands both the algorithms behind RL and the realities of distributed training and inference at scale.
- You'll wear many hats, from optimizing rollout and reward pipelines to enhancing reliability, observability, and orchestration, collaborating closely with researchers and infra teams to make reinforcement learning stable, fast, and production-ready.
- Note: This is an "evergreen role" that we keep open on an on-going basis to express interest.
Responsibilities
- Develop shared monitoring and observability tools to ensure high uptime, debuggability, and reproducibility for RL systems.
- Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines.
- Build evaluation and benchmarking infrastructure that measures model progress on helpfulness, safety, and factuality.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
Requirements
- Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases
Nice to have
- we encourage you to apply if you meet some but not all of these:
- Experience training or supporting large-scale language models with tens of billions of parameters or more.
- Background in high-performance or reliability engineering - distributed training frameworks and cluster orchestration (Kubernetes, Slurm).
- Familiarity with monitoring and observability tools (Prometheus, Grafana, OpenTelemetry).
- Contributions to large-scale ML research or infrastructure, open-source frameworks, or internal performance optimization efforts.
Skills
- Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
Compensation
- Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
Benefits
- Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
- Design, build, and optimize the infrastructure that powers large-scale reinforcement learning and post-training workloads.
- Bachelor's degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
Visa & Work Authorization
- While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- We sponsor visas.
Apply directly at Thinking Machines Lab →Create a free account for alerts like thisView Thinking Machines Lab immigration profile
This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.