Dyna Robotics

Dyna Robotics

ML Infrastructure Engineer, Training

Redwood City, CA

Sponsorship not specifiedDetected 112 days ago
Node.jsDistributed SystemsAWSGCPKubernetesMachine LearningPyTorchRoboticsResearchProblem Solving

About the role

  • As a ML Training Infrastructure Engineer, you will architect and build the systems that turn our multi-cloud GPU fleet into a training engine our researchers love.

Responsibilities

  • Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters. You'll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.
  • Optimize Researcher Ergonomics: Build a research codebase and job scheduling system (Kubernetes/SLURM) that prioritizes fast iteration, automated retries, and seamless failure recovery.
  • High-Performance Data Handling: Design high-throughput pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals), ensuring dataloaders never starve the GPUs.
  • Production Inference: Build low-latency inference pipelines for real-time robot control. You'll apply quantization, distillation, and model compilation (TensorRT, Triton) to move models from the lab to the physical world.
  • you design, build, and operate systems end-to-end to unblock fast-moving research.
  • Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters.
  • You'll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.
  • Ownership Mindset: You don't just "deploy" code; you design, build, and operate systems end-to-end to unblock fast-moving research.
  • own training infrastructure end-to-end so that every GPU is busy, every run is reproducible, and every researcher's next experiment is one command away.
  • Architect and own the infrastructure for large-scale GPU clusters.

Requirements

  • Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate).
  • ML Systems Mastery: Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate).
  • Experience with Robotics Data Formats (MCAP, Protobuf) or multimodal models (VLAs).

Skills

  • Hands-on experience managing cloud GPU environments (GCP/AWS) and container orchestration (Kubernetes).

Company info

  • JOIN US TO SHAPE THE NEXT FRONTIER OF AI-DRIVEN ROBOTICS!
  • Dyna Robotics makes general-purpose robots powered by a proprietary embodied AI foundation model that generalizes and self-improves across varied environments with commercial-grade performance.
  • Dyna's robots have been deployed at customers across multiple industries.
  • Its frontier model has the top generalization and performance in the industry.
  • Dyna Robotics was founded by repeat founders Lindon Gao and York Yang, who sold Caper AI for $350 million, and former DeepMind research scientist Jason Ma.
  • The company has raised over $140M, backed by top investors, including CRV and First Round.
  • We're positioned to redefine the landscape of robotic automation.
  • At Dyna Robotics, we build technology for the real world, which requires a team as diverse as the environments our robots inhabit.

Equal opportunity

  • We are an equal opportunity employer committed to technical rigor and mutual respect.

This listing is sourced directly from Dyna Robotics's careers page and normalized into a canonical job model.