Variance
Research Engineer, Evals
San Francisco
Sponsorship not specifiedDetected 113 days ago
Machine LearningLLMsAgentic AIA/B TestingResearch
About the role
- We're a small, talent-dense team in San Francisco working on a problem at the edge of what AI systems can reliably do: making good decisions in messy, adversarial, real-world environments.
- This role sits at the center of research, product, and engineering.
Responsibilities
- Build proprietary benchmarks and datasets to evaluate models and model systems on fraud, identity, and risk workflows
- Design and run offline and online evals that measure model performance on real customer tasks, not just abstract benchmarks
- Build reusable evaluation tools and quality building blocks that can be used across different product surfaces and workflows
- Partner closely with research, engineering, product, and design to improve system quality through rigorous experimentation
- Help create a strong culture of scientific experimentation, clear measurement, and continuous iteration
- We have a clear, trusted view of how our systems perform across the workflows that matter most
- We develop differentiated datasets, benchmarks, and quality loops that compound over time
- Experience building benchmarks, datasets, evaluation pipelines, or quality systems
- Ability to design clean experiments and draw reliable conclusions from noisy results
- Strong engineering judgment and a bias toward building
Nice to have
- Preferred background
Compensation
- Competitive salary and meaningful equity
Benefits
- Competitive salary and meaningful equity
- Platinum-level medical, dental, and vision insurance
- Unlimited PTO, sick leave, and parental leave
- Up to $100 per month in reimbursement for personal health and wellness expenses
Company info
- At Variance, we are teaching machines to make the hardest judgment calls at scale.
Apply directly at Variance →Create a free account for alerts like thisView Variance immigration profile
This listing is sourced directly from Variance's careers page and normalized into a canonical job model.