Ekarobotics
Machine Learning / Reinforcement Learning Infrastructure Engineer
Boston Area
Sponsorship not specifiedDetected 501 days ago
Distributed SystemsAWSGCPCloud PlatformsKubernetesCI/CDDevOpsMachine LearningDeep LearningPyTorchData EngineeringRoboticsTest AutomationResearchCommunication
About the role
- We are looking for a Reinforcement/Machine Learning Infrastructure Engineer to shape our training infrastructure.
- In this role, you will be responsible for designing, implementing, and maintaining the large-scale model training systems that power our next generation of robot learning.
- We believe that world-class infrastructure is the foundation for moving research into production.
Responsibilities
- Own Training Infrastructure: Design, implement, and maintain robust systems for large-scale model training, including job orchestration, scheduling, checkpointing, and experiment tracking.
- Developer Experience & Tooling: Build streamlined, intuitive abstractions for launching, monitoring, debugging, and reproducing experiments, minimizing friction and maximizing productivity for our research teams.
- Resource Management: Ensure efficient allocation and utilization of cloud-based compute resources while building the foundational systems needed for future scaling.
- Collaborate with Researchers: Partner with the research team to understand their needs, build infrastructure that supports cutting-edge methods, guide best practices for training at scale, and contribute to core JAX model and training code.
- You will focus on building an exceptional developer experience, creating intuitive and efficient tooling that our engineers and scientists love to use.
- Design, implement, and maintain robust systems for large-scale model training, including job orchestration, scheduling, checkpointing, and experiment tracking.
- Build streamlined, intuitive abstractions for launching, monitoring, debugging, and reproducing experiments, minimizing friction and maximizing productivity for our research teams.
- Partner with the research team to understand their needs, build infrastructure that supports cutting-edge methods, guide best practices for training at scale, and contribute to core JAX model and training code.
Requirements
- Familiarity with distributed training, multi-host setups, data pipelines, and managing workloads on cloud platforms or orchestration systems (e.g., Kubernetes, SLURM, GCP, AWS).
Nice to have
- Hands-on experience with large-scale training using JAX (preferred), PyTorch, or TensorFlow.
Benefits
- Scale Distributed Training: Work closely with researchers to reliably scale reinforcement learning and machine learning pipelines across compute clusters.
- Our team consists of pioneers in robotics and machine learning.
- This is a hands-on, high-impact role at the intersection of machine learning, software engineering, and scalable infrastructure.
Company info
- Eka Robotics is on a mission to build intelligence for the physical world - robots that are fast, general, and reliable.
- We are defining the frontier of robotics research and deployment.
- We are now hiring to scale our R&D effort.
- We are looking for hands-on individuals who are excited to help shape the future of robotics.
Apply directly at Ekarobotics →Create a free account for alerts like thisView Ekarobotics immigration profile
This listing is sourced directly from Ekarobotics's careers page and normalized into a canonical job model.