Inception
Member of Technical Staff, RL Infra
San Mateo, USA · Staff+
Sponsorship not specifiedDetected 134 days ago
DockerKubernetesCI/CDMachine LearningTensorFlowPyTorchAirflowAI OrchestrationSystems EngineeringResearch
About the role
- We're looking for engineers and scientists to design, optimize, and maintain the core systems that enable scalable, efficient reinforcement learning for large models.
- This role sits at the intersection of research and large-scale systems engineering: you'll wear many hats, from optimizing rollout and reward pipelines to enhancing reliability, observability, and orchestration, collaborating closely with researchers to make RL stable, fast, and production-ready.
Responsibilities
- Develop shared monitoring and observability tools to ensure high uptime, debuggability, and reproducibility for RL systems.
Nice to have
- Improve the reliability and scalability of RL training pipelines, distributed RL workloads, and training throughput.
- BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).
- Understanding of ML frameworks (PyTorch, TensorFlow, Ray, Megatron) from a systems perspective.
- Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines.
- Experience with ML workflow orchestration tools (Kubeflow, Airflow).
- Background in performance optimization and profiling of ML systems.
Benefits
- Design, build, and optimize the infrastructure that powers large-scale reinforcement learning and post-training workloads.
- Experience working with reinforcement learning workloads (PPO, DPO, RLHF, or reward modeling).
Apply directly at Inception →Create a free account for alerts like thisView Inception immigration profile
This listing is sourced directly from Inception's careers page and normalized into a canonical job model.