Mirendil

Mirendil

Member of Technical Staff, Post-Training, RL Infra

San Francisco · Staff+

Sponsorship not specified$300k-$500kDetected 28 days ago
AlgorithmsResearch

About the role

  • This role sits at the intersection of research and infrastructure.
  • You will work to push the scale of our RL stack, whether it is novel recipe ideas, reliability, or performance.
  • Some example areas you might work on (not limited to):

Responsibilities

  • Design and build reliable infrastructure for large-scale RL training
  • Implement novel performance optimizations across the training stack
  • Develop evaluation and benchmarking infrastructure to measure model progress, throughput, and uptime
  • Build data collection and feedback pipelines that close the loop between human signal, reward modeling, and training
  • Collaborate with multiple teams to rapidly iterate on RL algorithms and get experiments into production training runs

Compensation

  • We offer a base salary of $300,000–$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.

Benefits

  • We offer a base salary of $300,000-$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
  • Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments.

Company info

  • We are building a frontier AI research company and training our own models end-to-end.
  • We are looking for engineers to help build the post-training stack for frontier reasoning models.

This listing is sourced directly from Mirendil's careers page and normalized into a canonical job model.