Periodic Labs

Periodic Labs

ML Systems Engineer

Menlo Park

H1B sponsorship available$250k-$350kDetected 84 days ago
Machine LearningLLMsAgentic AILLMOpsSystems EngineeringCommunication

About the role

  • You'll work alongside some of the world's leading ML systems engineers, including leaders behind Megatron-LM, SGLang, Liger Kernel, TorchRec, CleanRL, TorchRL, and JAX-MD.

Responsibilities

  • Develop high-performance inference and serving systems
  • Design distributed runtimes and scheduling systems for complex ML workloads
  • Build secure and large-scale sandboxing and execution environments
  • Optimize memory, GPU kernels and communication for maximum throughput and end-to-end efficiency
  • Experience building high-performance ML infrastructure at scale
  • Ability to own complex technical problems end-to-end
  • Strong coding ability and engineering judgment, including the ability to work effectively with AI agents to design, implement, test, and debug complex systems
  • We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond.

Requirements

  • Sandboxing: Strong experience with secure execution environments, containers, virtualization, or code sandboxing.
  • GPU Kernels: Strong experience with CUDA, Triton, CUTLASS, CuTe, or custom GPU kernel development.
  • GPU Communication: Strong experience with NCCL, NVLink, InfiniBand, RDMA, GPUDirect RDMA, or large-scale communication optimization.
  • Strong experience with secure execution environments, containers, virtualization, or code sandboxing.
  • Strong experience with CUDA, Triton, CUTLASS, CuTe, or custom GPU kernel development.
  • Strong experience with NCCL, NVLink, InfiniBand, RDMA, GPUDirect RDMA, or large-scale communication optimization.
  • Bachelor's degree or similar experience

Nice to have

  • Familiarity with TorchTitan, FSDP, veRL, Slime, or other distributed training systems is a plus.
  • Strong experience with Ray.
  • Familiarity with Monarch or other distributed execution frameworks is a plus.
  • Strong experience with SGLang.
  • Familiarity with vLLM, TensorRT-LLM, or production LLM serving systems is a plus.
  • Distributed Runtime: Strong experience with Ray.
  • Inference: Strong experience with SGLang.

Compensation

  • $250,000-$350,000 base + equity

Benefits

  • Build and optimize large-scale training and reinforcement learning infrastructure while ensuring its correctness

Visa & Work Authorization

  • Yes, we sponsor visas and will do everything we can to assist in this process with our legal support.

This listing is sourced directly from Periodic Labs's careers page and normalized into a canonical job model.