Mirendil
Member of Technical Staff, Post-Training, RL
San Francisco · Staff+
Sponsorship not specified$300k-$500kDetected 28 days ago
Research
About the role
- This role sits at the point where model capability, training dynamics, data, verification, and infrastructure all meet.
- The best work here will involve both: forming hypotheses about training behavior, implementing them in real systems, running large-scale experiments, reading the resulting traces carefully, and turning the lessons into the next training run.
Responsibilities
- Develop and iterate on RL, SFT, and distillation recipes.
- Develop methods for assigning useful feedback across long trajectories, where sparse rewards, credit assignment, exploration, and verification all become harder.
- Build intuition and tooling for when off-policy data helps, when it hurts, and how to control the resulting instabilities.
- Build robust verification pipelines for tasks where correctness can be checked automatically or semi-automatically.
- Study the tradeoffs between specialization and generality, and design training mixtures that improve all capabilities together.
- Develop a deep empirical understanding of training runs.
- Diagnose regressions, separate real improvements from noise, design better ablations, and build the probes and analyses needed to make post-training less opaque.
- Verification and reward quality: Build robust verification pipelines for tasks where correctness can be checked automatically or semi-automatically.
- If you're excited about building the infrastructure that makes frontier RL research possible at scale, we'd love to hear from you.
Skills
- Mirendil is a tech-first company focused on solving core bottlenecks that unlock step-change acceleration across science and technology.
- Our first goal is to democratize frontier AI R&D across scientific disciplines.
- Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments.
- Researchers are also expected to have strong engineering skills.
Compensation
- We offer a base salary of $300,000–$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
Benefits
- We offer a base salary of $300,000-$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
Company info
- We are looking for research engineers to help build the post-training stack for frontier reasoning models.
Apply directly at Mirendil →Create a free account for alerts like thisView Mirendil immigration profile
This listing is sourced directly from Mirendil's careers page and normalized into a canonical job model.