Rhoda AI
Research Member of Technical Staff- Training Systems
Mountain View · Staff+
Sponsorship not specifiedDetected 65 days ago
Site Reliability EngineeringMachine LearningPyTorchCadenceSystems EngineeringRoboticsHardware DesignResearchCommunication
About the role
- We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.
- You will define how our models train at scale - driving efficiency, scalability, and correctness across large-scale multimodal training.
- Your work directly determines how efficiently we use compute, how well models scale across thousands of GPUs, and how quickly research can iterate.
Responsibilities
- Own training performance end-to-end
- Build systematic performance attribution: step-time decomposition (compute vs communication vs input pipeline), scaling curves across cluster sizes, and bottleneck identification and prioritization
- Drive measurable gains in:
- Design training systems (not just tune them)
- Build tools to identify bottlenecks quickly, track performance across model families, and compare scaling behavior across configurations
- Develop regression detection: microbenchmarks, performance baselines, and automated detection of efficiency regressions
- Partner deeply with researchers
- Collaborate on cluster-level efficiency
- We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots.
- Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design.
Requirements
- Proven track record improving large-scale distributed training performance
- Strong understanding of data / tensor / pipeline parallelism, sharded training (FSDP / ZeRO-style), communication patterns and overlap strategies, and scaling behavior across large GPU clusters
- Strong systems intuition. ability to reason across compute, communication, and memory bottlenecks
Nice to have
- High ownership mindset and comfort in a fast-moving environment
- GPU kernel or compiler-level experience (CUDA, Triton, graph capture, operator fusion)
- Experience with multimodal or video training (variable-length sequences, packing/bucketing)
- Experience working on large-scale training frameworks or distributed runtimes
- Familiarity with cluster topology, networking, and large-scale scheduling effects
- Direct leverage on research velocity - every efficiency gain you make accelerates model iteration across the entire research team
- Improvements you make compound across every training run the company executes - high ownership, high impact, small elite team
- Deep hands-on experience with modern ML stacks (PyTorch required
Skills
- Distributed efficiency (comm/compute overlap, bucketization, topology-aware mapping, parallelism strategies)
- Compute efficiency (kernel hotspots, operator fusion, attention optimization, framework/runtime overhead)
- Memory efficiency (activation checkpointing, sequence packing/bucketing, fragmentation reduction)
- Contribute to and extend training frameworks where needed
Benefits
- Diagnose and improve performance of large-scale multimodal training (vision, video, proprioception, actions, language)
Company info
- At Rhoda AI, we're building the next generation of generalist intelligent robots.
- We're looking for a Staff / Principal ML Training Systems Engineer to own training systems performance end-to-end.
Apply directly at Rhoda AI →Create a free account for alerts like thisView Rhoda AI immigration profile
This listing is sourced directly from Rhoda AI's careers page and normalized into a canonical job model.