Physical Intelligence
ML Infra Engineer (TPU/Jax/Optimization)
San Francisco
Sponsorship not specifiedDetected 179 days ago
AWSGCPCloud PlatformsKubernetesMachine LearningRoboticsResearchCommunication
About the role
- You'll work closely with researchers and model engineers to translate ideas into experiments-and those experiments into production training runs.
- This is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure.
Responsibilities
- Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging.
- Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization.
- Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments.
- Manage compute resources: Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost.
- Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale.
- Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics.
- Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms.
- Ability to debug and optimize performance bottlenecks across the training stack.
- In this role you will help scale and optimize our training systems and core model code.
- You'll own critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines.
Nice to have
- Hands-on large-scale training experience in JAX (preferred), PyTorch.
Skills
- Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost.
Company info
- The ML Infrastructure team supports and accelerates PI's core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast.
- The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs.
Apply directly at Physical Intelligence →Create a free account for alerts like thisView Physical Intelligence immigration profile
This listing is sourced directly from Physical Intelligence's careers page and normalized into a canonical job model.