Physical Intelligence

Physical Intelligence

ML Infra Engineer

San Francisco

Sponsorship not specifiedDetected 696 days ago
AWSGCPCloud PlatformsKubernetesMachine LearningRoboticsResearchCommunication

About the role

  • You'll work closely with researchers and model engineers to translate ideas into experiments-and those experiments into production training runs.
  • This is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure.
  • Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Responsibilities

  • Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging.
  • Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization.
  • Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments.
  • Manage compute resources: Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost.
  • Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale.
  • Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics.
  • Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms.
  • Ability to debug and optimize performance bottlenecks across the training stack.
  • In this role you will help scale and optimize our training systems and core model code.
  • You'll own critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines.

Nice to have

  • Hands-on large-scale training experience in JAX (preferred), PyTorch.

Skills

  • Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost.

Benefits

  • Bonus Points If You Have

Company info

  • The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs.

This listing is sourced directly from Physical Intelligence's careers page and normalized into a canonical job model.