Causal
Member of Technical Staff — Training Infrastructure
San Francisco · Staff+
Sponsorship not specifiedDetected 266 days ago
Machine LearningDeep LearningPyTorchLLMsRoboticsResearchCommunicationProblem Solving
About the role
- Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.
- We look for infrastructure engineers who are excited to tackle unsolved problems.
- Training an LPM means scaling novel architectures over multimodal physical data - a problem where the playbooks from language and vision only partially apply.
Responsibilities
- Design, implement, and optimize distributed training systems that scale across thousands of GPUs
- Build reusable frameworks for checkpointing, fault tolerance, and reproducibility that stay robust under rapid research iteration
- Collaborate with researchers to bring novel model architectures from prototype to full scale
Requirements
- We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
- Demonstrated proficiency with distributed training frameworks and techniques (e.g. FSDP, DeepSpeed, Megatron, Pytorch, JAX/XLA) to train large foundation models
- Ability to profile and debug performance in complex codebases, from framework internals down to kernels and collectives
Nice to have
- contributions to open-source ML infrastructure (e.g. PyTorch, Megatron-LM, DeepSpeed, XLA)
Benefits
- Deep understanding of deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures
- Bonus: contributions to open-source ML infrastructure (e.g. PyTorch, Megatron-LM, DeepSpeed, XLA)
Company info
- What we're looking for
- Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it.
This listing is sourced directly from Causal's careers page and normalized into a canonical job model.