Inception

Inception

Member of Technical Staff, RL Infra

San Mateo, USA · Staff+

Sponsorship not specifiedDetected 134 days ago
DockerKubernetesCI/CDMachine LearningTensorFlowPyTorchAirflowAI OrchestrationSystems EngineeringResearch

About the role

  • We're looking for engineers and scientists to design, optimize, and maintain the core systems that enable scalable, efficient reinforcement learning for large models.
  • This role sits at the intersection of research and large-scale systems engineering: you'll wear many hats, from optimizing rollout and reward pipelines to enhancing reliability, observability, and orchestration, collaborating closely with researchers to make RL stable, fast, and production-ready.

Responsibilities

  • Develop shared monitoring and observability tools to ensure high uptime, debuggability, and reproducibility for RL systems.

Nice to have

  • Improve the reliability and scalability of RL training pipelines, distributed RL workloads, and training throughput.
  • BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).
  • Understanding of ML frameworks (PyTorch, TensorFlow, Ray, Megatron) from a systems perspective.
  • Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines.
  • Experience with ML workflow orchestration tools (Kubeflow, Airflow).
  • Background in performance optimization and profiling of ML systems.

Benefits

  • Design, build, and optimize the infrastructure that powers large-scale reinforcement learning and post-training workloads.
  • Experience working with reinforcement learning workloads (PPO, DPO, RLHF, or reward modeling).

This listing is sourced directly from Inception's careers page and normalized into a canonical job model.