Mirendil

Mirendil

Member of Technical Staff, Post-Training, RL

San Francisco · Staff+

Sponsorship not specified$300k-$500kDetected 28 days ago
Research

About the role

  • This role sits at the point where model capability, training dynamics, data, verification, and infrastructure all meet.
  • The best work here will involve both: forming hypotheses about training behavior, implementing them in real systems, running large-scale experiments, reading the resulting traces carefully, and turning the lessons into the next training run.

Responsibilities

  • Develop and iterate on RL, SFT, and distillation recipes.
  • Develop methods for assigning useful feedback across long trajectories, where sparse rewards, credit assignment, exploration, and verification all become harder.
  • Build intuition and tooling for when off-policy data helps, when it hurts, and how to control the resulting instabilities.
  • Build robust verification pipelines for tasks where correctness can be checked automatically or semi-automatically.
  • Study the tradeoffs between specialization and generality, and design training mixtures that improve all capabilities together.
  • Develop a deep empirical understanding of training runs.
  • Diagnose regressions, separate real improvements from noise, design better ablations, and build the probes and analyses needed to make post-training less opaque.
  • Verification and reward quality: Build robust verification pipelines for tasks where correctness can be checked automatically or semi-automatically.
  • If you're excited about building the infrastructure that makes frontier RL research possible at scale, we'd love to hear from you.

Skills

  • Mirendil is a tech-first company focused on solving core bottlenecks that unlock step-change acceleration across science and technology.
  • Our first goal is to democratize frontier AI R&D across scientific disciplines.
  • Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments.
  • Researchers are also expected to have strong engineering skills.

Compensation

  • We offer a base salary of $300,000–$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.

Benefits

  • We offer a base salary of $300,000-$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.

Company info

  • We are looking for research engineers to help build the post-training stack for frontier reasoning models.

This listing is sourced directly from Mirendil's careers page and normalized into a canonical job model.