Causal
Member of Technical Staff — Compute Cluster
San Francisco · Staff+
Sponsorship not specifiedDetected 4 days ago
AWSGCPAzureCloud PlatformsDockerKubernetesLinuxMachine LearningAI OrchestrationRoboticsResearchProblem Solving
About the role
- Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.
- We look for infrastructure engineers who are excited to tackle unsolved problems.
- Everything we do - training, evaluation, serving - runs on our GPU fleet.
Responsibilities
- Design, deploy, and operate large distributed GPU clusters end to end: provisioning, imaging, upgrades, and capacity planning
- Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers
- Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage
- build the observability to catch failures before researchers do
- Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs
- Monitor and continuously improve reliability and error recovery; build the observability to catch failures before researchers do
Requirements
- We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
- Experience operating large-scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker)
- Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings
- Familiarity with CUDA/NCCL and performance profiling for distributed workloads
- We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather.
Company info
- What we're looking for
This listing is sourced directly from Causal's careers page and normalized into a canonical job model.