Ekarobotics

Ekarobotics

Machine Learning / Reinforcement Learning Infrastructure Engineer

Boston Area

Sponsorship not specifiedDetected 501 days ago
Distributed SystemsAWSGCPCloud PlatformsKubernetesCI/CDDevOpsMachine LearningDeep LearningPyTorchData EngineeringRoboticsTest AutomationResearchCommunication

About the role

  • We are looking for a Reinforcement/Machine Learning Infrastructure Engineer to shape our training infrastructure.
  • In this role, you will be responsible for designing, implementing, and maintaining the large-scale model training systems that power our next generation of robot learning.
  • We believe that world-class infrastructure is the foundation for moving research into production.

Responsibilities

  • Own Training Infrastructure: Design, implement, and maintain robust systems for large-scale model training, including job orchestration, scheduling, checkpointing, and experiment tracking.
  • Developer Experience & Tooling: Build streamlined, intuitive abstractions for launching, monitoring, debugging, and reproducing experiments, minimizing friction and maximizing productivity for our research teams.
  • Resource Management: Ensure efficient allocation and utilization of cloud-based compute resources while building the foundational systems needed for future scaling.
  • Collaborate with Researchers: Partner with the research team to understand their needs, build infrastructure that supports cutting-edge methods, guide best practices for training at scale, and contribute to core JAX model and training code.
  • You will focus on building an exceptional developer experience, creating intuitive and efficient tooling that our engineers and scientists love to use.
  • Design, implement, and maintain robust systems for large-scale model training, including job orchestration, scheduling, checkpointing, and experiment tracking.
  • Build streamlined, intuitive abstractions for launching, monitoring, debugging, and reproducing experiments, minimizing friction and maximizing productivity for our research teams.
  • Partner with the research team to understand their needs, build infrastructure that supports cutting-edge methods, guide best practices for training at scale, and contribute to core JAX model and training code.

Requirements

  • Familiarity with distributed training, multi-host setups, data pipelines, and managing workloads on cloud platforms or orchestration systems (e.g., Kubernetes, SLURM, GCP, AWS).

Nice to have

  • Hands-on experience with large-scale training using JAX (preferred), PyTorch, or TensorFlow.

Benefits

  • Scale Distributed Training: Work closely with researchers to reliably scale reinforcement learning and machine learning pipelines across compute clusters.
  • Our team consists of pioneers in robotics and machine learning.
  • This is a hands-on, high-impact role at the intersection of machine learning, software engineering, and scalable infrastructure.

Company info

  • Eka Robotics is on a mission to build intelligence for the physical world - robots that are fast, general, and reliable.
  • We are defining the frontier of robotics research and deployment.
  • We are now hiring to scale our R&D effort.
  • We are looking for hands-on individuals who are excited to help shape the future of robotics.

This listing is sourced directly from Ekarobotics's careers page and normalized into a canonical job model.