Thinking Machines Lab

Thinking Machines Lab

Network Engineer, Supercomputing

San Francisco

H1B sponsorship available$350k-$475kDetected 27 days ago
PythonRustExpressKubernetesLinuxDeep LearningPyTorchLogisticsNetwork EngineeringCommunication

About the role

  • A single degraded link or flapping NIC can quietly slow a long training run or take it down outright
  • you'll be responsible for interconnect reliability at scale, across large GPU fabrics - both the RDMA/RoCE fabric between nodes and the NVLink/NVSwitch domains within them.
  • Your goal is for our researchers to trust the fleet without worrying about the fabric underneath.

Responsibilities

  • Reason about and validate GPU network fabric design across our deployments.
  • Build host-level network instrumentation and use Linux tooling to build dashboards and alerts, not just the bug report.
  • Drive escalations with cloud-provider networking teams, owning issues end-to-end until they're resolved.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (we use Python or Rust).
  • Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Extensive experience with at least one of the following:

Nice to have

  • we encourage you to apply if you meet some but not all of these:
  • Fluency with host-level debugging tools on Linux.
  • Strong communication skills, internally and with cloud providers.
  • Familiarity with cloud network primitives across at least two cloud providers.
  • Hands-on experience with NVLink / NVSwitch, fabric manager, and IMEX.
  • Statistical rigor in reliability reasoning - comfort reasoning about failure and error rates, distributions, and base rates, and the judgment to separate signal from noise when characterizing a large fabric.
  • A track record of writing tooling that made the next debugging session meaningfully faster.
  • Familiarity with CUDA/NCCL and performance profiling for distributed training and inference.

Skills

  • Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.

Compensation

  • Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

Benefits

  • Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
  • Own NVLink / NVSwitch interconnect - including fabric manager and IMEX health, link and lane errors, and how the GPU fabric interacts with collectives.

Visa & Work Authorization

  • While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • We sponsor visas.

This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.