Thinking Machines Lab

Thinking Machines Lab

Software Engineer, Systems Generalist

San Francisco

H1B sponsorship available$350k-$475kDetected 110 days ago
PythonRustExpressDistributed SystemsKubernetesMachine LearningPyTorchSparkData EngineeringLogisticsResearch

About the role

  • You'll join a small, high-impact team responsible for architecting and scaling the core infrastructure behind everything we do.
  • Infrastructure is critical to us: it's the bedrock that enables every breakthrough.
  • You'll work directly with researchers to accelerate experiments, improve infrastructure efficiency, and enable key insights across our models, products, and data assets.

Responsibilities

  • Core Infrastructure: We support teams that train, research, and ultimately serve AI models and build the underlying infrastructure for the clusters to reliably and safely train frontier models.
  • Examples might include building systems and running large Kubernetes clusters with GPU workloads, or building infrastructure to support Tinker.
  • Data Infrastructure: We build and maintain the data systems for our research and products.
  • You'll design and optimize data pipelines using tools like Spark and other modern data infrastructure technologies.
  • You'll build scalable, reliable, data infrastructure while embedding governance best practices.
  • Developer Productivity: We care deeply about research and engineering productivity and our ability to continue shipping quickly. We build tooling, systems, frameworks, and systems to make sure everyone gets well configured, optimized developer environments.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
  • We support teams that train, research, and ultimately serve AI models and build the underlying infrastructure for the clusters to reliably and safely train frontier models.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (we use Python or Rust).
  • Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • We care deeply about research and engineering productivity and our ability to continue shipping quickly.

Nice to have

  • we encourage you to apply if you meet some but not all of these:
  • Strong debugging across application, OS, and network layers.
  • Proficiency in Python or Rust (or similar), containers, and modern CI.
  • Experience with Kubernetes, controllers/operators, or performance profiling.
  • Familiarity with GPU/ML workflows or large‑scale data/eval pipelines.

Skills

  • Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.

Compensation

  • Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

Benefits

  • Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
  • This flexible approach allows us to match talented engineers with the infrastructure teams where they'll have the greatest impact and growth potential.

Visa & Work Authorization

  • While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • We sponsor visas.

This listing is sourced directly from Thinking Machines Lab's careers page and normalized into a canonical job model.