Causal

Causal

Member of Technical Staff — Compute Cluster

San Francisco · Staff+

Sponsorship not specifiedDetected 4 days ago
AWSGCPAzureCloud PlatformsDockerKubernetesLinuxMachine LearningAI OrchestrationRoboticsResearchProblem Solving

About the role

  • Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.
  • We look for infrastructure engineers who are excited to tackle unsolved problems.
  • Everything we do - training, evaluation, serving - runs on our GPU fleet.

Responsibilities

  • Design, deploy, and operate large distributed GPU clusters end to end: provisioning, imaging, upgrades, and capacity planning
  • Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers
  • Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage
  • build the observability to catch failures before researchers do
  • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs
  • Monitor and continuously improve reliability and error recovery; build the observability to catch failures before researchers do

Requirements

  • We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
  • Experience operating large-scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker)
  • Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings
  • Familiarity with CUDA/NCCL and performance profiling for distributed workloads
  • We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather.

Company info

  • What we're looking for

This listing is sourced directly from Causal's careers page and normalized into a canonical job model.