Thinking Machines Lab
Network Engineer, Supercomputing
San Francisco
H1B sponsorship available$350k-$475kDetected 27 days ago
PythonRustExpressKubernetesLinuxDeep LearningPyTorchLogisticsNetwork EngineeringCommunication
About the role
- A single degraded link or flapping NIC can quietly slow a long training run or take it down outright
- you'll be responsible for interconnect reliability at scale, across large GPU fabrics - both the RDMA/RoCE fabric between nodes and the NVLink/NVSwitch domains within them.
- Your goal is for our researchers to trust the fleet without worrying about the fabric underneath.
Responsibilities
- Reason about and validate GPU network fabric design across our deployments.
- Build host-level network instrumentation and use Linux tooling to build dashboards and alerts, not just the bug report.
- Drive escalations with cloud-provider networking teams, owning issues end-to-end until they're resolved.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, or similar.
- Proficiency in at least one backend language (we use Python or Rust).
- Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
- Comfort operating across the stack and owning projects end-to-end.
- Extensive experience with at least one of the following:
Nice to have
- we encourage you to apply if you meet some but not all of these:
- Fluency with host-level debugging tools on Linux.
- Strong communication skills, internally and with cloud providers.
- Familiarity with cloud network primitives across at least two cloud providers.
- Hands-on experience with NVLink / NVSwitch, fabric manager, and IMEX.
- Statistical rigor in reliability reasoning - comfort reasoning about failure and error rates, distributions, and base rates, and the judgment to separate signal from noise when characterizing a large fabric.
- A track record of writing tooling that made the next debugging session meaningfully faster.
- Familiarity with CUDA/NCCL and performance profiling for distributed training and inference.
Skills
- Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
Compensation
- Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
Benefits
- Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
- Own NVLink / NVSwitch interconnect - including fabric manager and IMEX health, link and lane errors, and how the GPU fabric interacts with collectives.
Visa & Work Authorization
- While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- We sponsor visas.
Apply directly at Thinking Machines Lab →Create a free account for alerts like thisView Thinking Machines Lab immigration profile
This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.