Thinking Machines Lab
Reliability Engineer, Supercomputing
San Francisco
H1B sponsorship available$350k-$475kDetected 27 days ago
PythonRustExpressKubernetesLinuxPyTorchLogisticsResearchWriting
About the role
- We're hiring an engineer to ensure the reliability of our GPU supercomputing fleet, owning the seam between hardware, firmware, and operating system.
- You will track the long tail of hardware issues: We are conducting frontier research in AI and a single bad NIC, HBM or a kernel driver edge case can compromise an experiment.
- Your job is to diagnose these issues, track their root cause down to the hardware, and resolve them internally or directly with vendors so that our researchers can run at scale and with confidence.
Responsibilities
- Own the drivers, kernel surface, and diagnostics that span hardware, firmware, and OS.
- Automate the monitoring of fleet reliability and analyze error rates to validate whether a fix or firmware change measurably reduced failures rather than shifting them around.
- Drive the firmware lifecycle: tracking, qualification, staged rollout, and regression analysis.
- Engage vendors directly - GPUs, server OEMs, NIC vendors, and storage vendors - to get real fixes rather than ticket numbers. Manage RMA flows when hardware needs to come out.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
- Manage RMA flows when hardware needs to come out.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, or similar.
- Proficiency in at least one backend language (we use Python or Rust).
- Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
- Comfort operating across the stack and owning projects end-to-end.
- Proficiency in at least one backend language (we use Python and Rust).
Nice to have
- we encourage you to apply if you meet some but not all of these:
- Fluency with Linux systems and debugging tools.
- Proven statistical rigor in analyzing reliability.
- A track record of debugging a problem from application symptom to the root cause in hardware.
- Comfort reading vendor errata, firmware release notes, and kernel changelogs.
- Experience engaging hardware vendors directly - not just through escalation portals.
Skills
- Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
Compensation
- Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
Benefits
- Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
- Monitor and improve GPU hardware health signals and turn them into actionable reliability improvements.
Company info
- We are conducting frontier research in AI and a single bad NIC, HBM or a kernel driver edge case can compromise an experiment.
Visa & Work Authorization
- While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- We sponsor visas.
Apply directly at Thinking Machines Lab →Create a free account for alerts like thisView Thinking Machines Lab immigration profile
This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.