Mirendil
Member of Technical Staff, Model Evaluation
San Francisco · Staff+
Sponsorship not specified$300k-$500kDetected 28 days ago
Detection EngineeringResearch
About the role
- Some example areas you might work on (not limited to):
- Instrument training runs with observability tooling so researchers can understand what's changing in model behavior, and why.
- If you're excited about the hard problem of knowing whether a frontier AI system is actually improving, we'd love to hear from you.
Responsibilities
- Build automated eval pipelines and regression-detection systems that run continuously and surface signal quickly.
- Develop agent-assisted workflows for humans to efficiently inspect model behavior.
- Partner with post-training and RL teams to close the loop between eval signal and training decisions.
- Design and build evaluation frameworks that measure model capabilities along realistic axes, beyond standard benchmarks.
Skills
- Mirendil is a tech-first company focused on solving core bottlenecks that unlock step-change acceleration across science and technology.
- Our first goal is to democratize frontier AI R&D across scientific disciplines.
- Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments.
Compensation
- We offer a base salary of $300,000–$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
Benefits
- We offer a base salary of $300,000-$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
Company info
- We are looking for a research engineer to build the evaluation infrastructure that tells us whether our models are getting better in ways we care about.
Apply directly at Mirendil →Create a free account for alerts like thisView Mirendil immigration profile
This listing is sourced directly from Mirendil's careers page and normalized into a canonical job model.