Simile
Evaluations Engineering - Member of Technical Staff
San Francisco · Staff+
Sponsorship not specified$200k-$400kDetected 5 days ago
Backend DevelopmentMachine LearningData EngineeringData ScienceLLMsStatisticsResearchCommunication
About the role
- You will work across data and evaluation infrastructure, evaluation execution workflows, backend services, automation, and internal tooling.
- Your initial focus will include streamlining how evaluations are run across models
- strengthening evaluation versioning, data models, and access controls
Responsibilities
- Develop the services, pipelines, and orchestration needed to run evaluations efficiently across datasets, model versions, populations, and use cases.
- Design relational schemas, versioning, provenance, permissions, and quality controls that make evaluation results reproducible and trustworthy.
- Partner with Evals and Data Operations to streamline customer validations, survey deployment, response ingestion, and the integration of new ground truth.
- Build interfaces that help teams manage evals, compare models, investigate results, and identify regressions.
- Our hiring journey is designed to help both sides align on fit, working style, and expectations.
- As a Member of Technical Staff in Evaluations Engineering, you will build the systems that enable Simile to evaluate whether our simulations of human behavior are accurate, trustworthy, and improving over time.
- Build evaluation execution infrastructure: Develop the services, pipelines, and orchestration needed to run evaluations efficiently across datasets, model versions, populations, and use cases.
- Strengthen evaluation data systems: Design relational schemas, versioning, provenance, permissions, and quality controls that make evaluation results reproducible and trustworthy.
- Automate validation and data collection: Partner with Evals and Data Operations to streamline customer validations, survey deployment, response ingestion, and the integration of new ground truth.
- Develop evaluation tooling: Build interfaces that help teams manage evals, compare models, investigate results, and identify regressions.
Requirements
- Strong Engineering Fundamentals: Several years of experience building and maintaining production-quality software, with sound judgment in system design, testing, debugging, and maintainability.
- Ownership and Communication: A track record of independently driving important technical work and collaborating effectively across engineering, research, and operations.
- Human Data Systems: Experience with labeling platforms, expert-review workflows, LLM-as-judge systems, grader calibration, or other human-in-the-loop evaluation methods.
- Sensitive Data and Access Controls: Experience designing permissions, auditability, and data-governance systems for human or customer data.
- Agentic Engineering: Experience using modern AI coding tools to accelerate development while independently testing and validating their output.
- Several years of experience building and maintaining production-quality software, with sound judgment in system design, testing, debugging, and maintainability.
- Experience with labeling platforms, expert-review workflows, LLM-as-judge systems, grader calibration, or other human-in-the-loop evaluation methods.
- Experience using modern AI coding tools to accelerate development while independently testing and validating their output.
Skills
- Final offers are based on experience, specialized skills, interview performance, and relevant training.
Compensation
- $200,000 - $400,000 USD
- At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.
Benefits
- At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.
- Comprehensive medical, dental, and vision coverage.
- Flexible time off policies to support work-life balance.
- Equity: Grants are available for eligible roles, subject to board approval.
- application process due to a disability, please let us know.
Company info
- Pilots don't train with real passengers.
- Actors don't rehearse with real audiences.
- Yet, the most consequential decisions in society are often pushed straight to production.
- Simile is changing that.
- We have built the first AI simulation of society, populated by generative agents based on real humans.
- Our research pioneered the field of AI-based simulation, proving it is possible to model human behavior with high accuracy.
- Today, we are developing a Foundation Model to predict human behavior in any situation, at any scale.
- We are backed by $100M in funding led by Index Ventures, with participation from Hanabi, A*, Bain Capital Ventures, and AI visionaries including Andrej Karpathy, Fei-Fei Li, Adam D'Angelo, and Guillermo Rauch.
- A track record of independently driving important technical work and collaborating effectively across engineering, research, and operations.
- We do not expect one person to have all of these.
- We are hiring a team with complementary strengths.
- We are happy to assist.
Equal opportunity
- Simile is an equal opportunity workplace.
- We welcome applicants of all backgrounds and identities, valuing an environment where everyone can contribute authentically.
- Accommodations: If you require support or reasonable accommodations during the application process due to a disability, please let us know.
This listing is sourced directly from Simile's careers page and normalized into a canonical job model.