Physical Intelligence

Physical Intelligence

ML Infra Engineer (TPU/Jax/Optimization)

San Francisco

Sponsorship not specifiedDetected 179 days ago
AWSGCPCloud PlatformsKubernetesMachine LearningRoboticsResearchCommunication

About the role

  • You'll work closely with researchers and model engineers to translate ideas into experiments-and those experiments into production training runs.
  • This is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure.

Responsibilities

  • Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging.
  • Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization.
  • Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments.
  • Manage compute resources: Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost.
  • Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale.
  • Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics.
  • Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms.
  • Ability to debug and optimize performance bottlenecks across the training stack.
  • In this role you will help scale and optimize our training systems and core model code.
  • You'll own critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines.

Nice to have

  • Hands-on large-scale training experience in JAX (preferred), PyTorch.

Skills

  • Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost.

Company info

  • The ML Infrastructure team supports and accelerates PI's core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast.
  • The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs.

This listing is sourced directly from Physical Intelligence's careers page and normalized into a canonical job model.