Mirendil

Mirendil

Member of Technical Staff, Model Evaluation

San Francisco · Staff+

Sponsorship not specified$300k-$500kDetected 28 days ago
Detection EngineeringResearch

About the role

  • Some example areas you might work on (not limited to):
  • Instrument training runs with observability tooling so researchers can understand what's changing in model behavior, and why.
  • If you're excited about the hard problem of knowing whether a frontier AI system is actually improving, we'd love to hear from you.

Responsibilities

  • Build automated eval pipelines and regression-detection systems that run continuously and surface signal quickly.
  • Develop agent-assisted workflows for humans to efficiently inspect model behavior.
  • Partner with post-training and RL teams to close the loop between eval signal and training decisions.
  • Design and build evaluation frameworks that measure model capabilities along realistic axes, beyond standard benchmarks.

Skills

  • Mirendil is a tech-first company focused on solving core bottlenecks that unlock step-change acceleration across science and technology.
  • Our first goal is to democratize frontier AI R&D across scientific disciplines.
  • Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments.

Compensation

  • We offer a base salary of $300,000–$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.

Benefits

  • We offer a base salary of $300,000-$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.

Company info

  • We are looking for a research engineer to build the evaluation infrastructure that tells us whether our models are getting better in ways we care about.

This listing is sourced directly from Mirendil's careers page and normalized into a canonical job model.