Rhoda AI

Rhoda AI

Research Member of Technical Staff- Training Systems

Mountain View · Staff+

Sponsorship not specifiedDetected 65 days ago
Site Reliability EngineeringMachine LearningPyTorchCadenceSystems EngineeringRoboticsHardware DesignResearchCommunication

About the role

  • We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.
  • You will define how our models train at scale - driving efficiency, scalability, and correctness across large-scale multimodal training.
  • Your work directly determines how efficiently we use compute, how well models scale across thousands of GPUs, and how quickly research can iterate.

Responsibilities

  • Own training performance end-to-end
  • Build systematic performance attribution: step-time decomposition (compute vs communication vs input pipeline), scaling curves across cluster sizes, and bottleneck identification and prioritization
  • Drive measurable gains in:
  • Design training systems (not just tune them)
  • Build tools to identify bottlenecks quickly, track performance across model families, and compare scaling behavior across configurations
  • Develop regression detection: microbenchmarks, performance baselines, and automated detection of efficiency regressions
  • Partner deeply with researchers
  • Collaborate on cluster-level efficiency
  • We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots.
  • Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design.

Requirements

  • Proven track record improving large-scale distributed training performance
  • Strong understanding of data / tensor / pipeline parallelism, sharded training (FSDP / ZeRO-style), communication patterns and overlap strategies, and scaling behavior across large GPU clusters
  • Strong systems intuition. ability to reason across compute, communication, and memory bottlenecks

Nice to have

  • High ownership mindset and comfort in a fast-moving environment
  • GPU kernel or compiler-level experience (CUDA, Triton, graph capture, operator fusion)
  • Experience with multimodal or video training (variable-length sequences, packing/bucketing)
  • Experience working on large-scale training frameworks or distributed runtimes
  • Familiarity with cluster topology, networking, and large-scale scheduling effects
  • Direct leverage on research velocity - every efficiency gain you make accelerates model iteration across the entire research team
  • Improvements you make compound across every training run the company executes - high ownership, high impact, small elite team
  • Deep hands-on experience with modern ML stacks (PyTorch required

Skills

  • Distributed efficiency (comm/compute overlap, bucketization, topology-aware mapping, parallelism strategies)
  • Compute efficiency (kernel hotspots, operator fusion, attention optimization, framework/runtime overhead)
  • Memory efficiency (activation checkpointing, sequence packing/bucketing, fragmentation reduction)
  • Contribute to and extend training frameworks where needed

Benefits

  • Diagnose and improve performance of large-scale multimodal training (vision, video, proprioception, actions, language)

Company info

  • At Rhoda AI, we're building the next generation of generalist intelligent robots.
  • We're looking for a Staff / Principal ML Training Systems Engineer to own training systems performance end-to-end.

This listing is sourced directly from Rhoda AI's careers page and normalized into a canonical job model.