Dyna Robotics
ML Infrastructure Engineer, Training
Redwood City, CA
Sponsorship not specifiedDetected 112 days ago
Node.jsDistributed SystemsAWSGCPKubernetesMachine LearningPyTorchRoboticsResearchProblem Solving
About the role
- As a ML Training Infrastructure Engineer, you will architect and build the systems that turn our multi-cloud GPU fleet into a training engine our researchers love.
Responsibilities
- Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters. You'll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.
- Optimize Researcher Ergonomics: Build a research codebase and job scheduling system (Kubernetes/SLURM) that prioritizes fast iteration, automated retries, and seamless failure recovery.
- High-Performance Data Handling: Design high-throughput pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals), ensuring dataloaders never starve the GPUs.
- Production Inference: Build low-latency inference pipelines for real-time robot control. You'll apply quantization, distillation, and model compilation (TensorRT, Triton) to move models from the lab to the physical world.
- you design, build, and operate systems end-to-end to unblock fast-moving research.
- Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters.
- You'll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.
- Ownership Mindset: You don't just "deploy" code; you design, build, and operate systems end-to-end to unblock fast-moving research.
- own training infrastructure end-to-end so that every GPU is busy, every run is reproducible, and every researcher's next experiment is one command away.
- Architect and own the infrastructure for large-scale GPU clusters.
Requirements
- Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate).
- ML Systems Mastery: Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate).
- Experience with Robotics Data Formats (MCAP, Protobuf) or multimodal models (VLAs).
Skills
- Hands-on experience managing cloud GPU environments (GCP/AWS) and container orchestration (Kubernetes).
Company info
- JOIN US TO SHAPE THE NEXT FRONTIER OF AI-DRIVEN ROBOTICS!
- Dyna Robotics makes general-purpose robots powered by a proprietary embodied AI foundation model that generalizes and self-improves across varied environments with commercial-grade performance.
- Dyna's robots have been deployed at customers across multiple industries.
- Its frontier model has the top generalization and performance in the industry.
- Dyna Robotics was founded by repeat founders Lindon Gao and York Yang, who sold Caper AI for $350 million, and former DeepMind research scientist Jason Ma.
- The company has raised over $140M, backed by top investors, including CRV and First Round.
- We're positioned to redefine the landscape of robotic automation.
- At Dyna Robotics, we build technology for the real world, which requires a team as diverse as the environments our robots inhabit.
Equal opportunity
- We are an equal opportunity employer committed to technical rigor and mutual respect.
Apply directly at Dyna Robotics →Create a free account for alerts like thisView Dyna Robotics immigration profile
This listing is sourced directly from Dyna Robotics's careers page and normalized into a canonical job model.