Rhoda AI
Senior Inference Optimization ML Engineer
Mountain View · Senior
Sponsorship not specifiedDetected 71 days ago
Machine LearningPyTorchLLMsRoboticsHardware DesignResearch
About the role
- We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.
- You'll be responsible for squeezing maximum performance out of large multimodal models, across cloud and on-robot deployment targets.
- You will working closely with research and robotics teams to close the gap between training and real-world deployment.
Responsibilities
- Own inference performance end-to-end - diagnose and improve latency, throughput, and efficiency of large foundation models in production
- Build systematic performance attribution: latency decomposition (compute vs. memory bandwidth vs. I/O), bottleneck identification, and prioritization across model families
- Apply and develop optimization techniques including quantization, pruning, distillation, operator fusion, and model compilation (e.g., TensorRT, torch.compile, XLA)
- Work with kernel-level tooling (e.g., CUDA, Triton) to identify hotspots and implement or tune custom kernels where needed
- Build benchmarking and regression detection infrastructure: latency baselines, throughput curves, and automated detection of performance regressions across model versions
- Collaborate closely with research engineers to translate model innovations into optimized, deployment-ready implementations
Requirements
- 3+ years of experience in inference optimization, ML systems, or a closely related field
- Strong understanding of compute, memory bandwidth, and I/O bottlenecks in large model inference
- Experience with model optimization techniques: quantization (INT8/FP8/AWQ), distillation, pruning, and compilation
- Familiarity with inference serving frameworks (e.g., Triton, TensorRT, vLLM, TorchServe)
Nice to have
- High ownership mindset and comfort in a fast-moving environment
- GPU kernel or compiler-level experience (CUDA, Triton, graph capture, operator fusion)
- Experience with multimodal or video model inference (variable-length sequences, packing/bucketing)
- Experience with speculative decoding, continuous batching, or other LLM serving optimizations
- Background in streaming or low-latency systems relevant to real-time robot control
- Direct leverage on research velocity and real-world robot performance - every efficiency gain you make accelerates model iteration and tightens the loop between model and robot behavior
- Deep hands-on experience with modern ML stacks (PyTorch required
- Nice to Have (But Not Required)
Benefits
- Optimize attention mechanisms, KV caching, and memory layouts for large multimodal models (vision, video, language, proprioception)
Company info
- What We're Looking For
Apply directly at Rhoda AI →Create a free account for alerts like thisView Rhoda AI immigration profile
This listing is sourced directly from Rhoda AI's careers page and normalized into a canonical job model.