Simile
Evaluations - Member of Technical Staff
San Francisco · Staff+
Sponsorship not specified$200k-$400kDetected 21 days ago
PythonSQLMachine LearningData ScienceLLMsStatisticsA/B TestingResearchExperimental Design
About the role
- You will help shape what Simile measures, the quality bars we defend, and how evaluation evidence guides model, product, and customer decisions.
- Evaluation at Simile brings together model evals, statistics, behavioral science, research methodology, product quality, and human judgment.
- Our models simulate people, populations, markets, and groups, which means our evals must reason about distributions, noisy human ground truth, uncertainty, qualitative outputs, behavioral data, and customer decision-making.
Responsibilities
- Design evals, metrics, rubrics, datasets, dashboards, and workflows that measure whether Simile's models are accurately predicting human behavior across customer use cases, populations, question types, and decision contexts.
- Build evals for qualitative responses, retrieval, survey generation, AI-generated research reports, customer-facing outputs, and other product surfaces where model quality directly shapes customer trust.
- Develop rigorous ways to compare simulated responses against human data, customer studies, Simile-collected ground truth, and behavioral datasets.
- Experience building model evaluation dashboards, regression suites, release gates, benchmark sets, model comparison workflows, or systems that help ML teams decide where to focus and when to ship.
- Our hiring journey is designed to help both sides align on fit, working style, and expectations.
Requirements
- Evaluation Taste: You have strong intuition for what makes an eval meaningful, robust, and decision-relevant.
- You can explain what an eval measures, what it does not measure, how it can be gamed, and why it should or should not affect a model or product decision.
- You do not need to be a modeling specialist, but you can read model outputs, understand modeling team needs, and reason about whether a model change actually improved the thing we care about.
- You are comfortable working with data and automation tools such as Python, SQL, R, notebooks, LLM APIs, and agentic coding tools such as Codex, Claude Code, Cursor, or equivalent systems.
- You know how to move quickly while still validating outputs, catching errors, and planning for the long-term..
- Survey Methodology and Statistics: Experience with sampling, weighting, margin of error, power analysis, uncertainty quantification, Bayesian modeling, causal inference, psychometrics, polling, or measurement theory.
- Multi-Agent or Group Behavior: Interest or experience in modeling group conversation, deliberation, focus groups, juries, committees, polarization, collective decision-making, or social influence.
Skills
- Final offers are based on experience, specialized skills, interview performance, and relevant training.
Compensation
- $200,000 - $400,000 USD
- At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.
Benefits
- At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.
- Comprehensive medical, dental, and vision coverage.
- Flexible time off policies to support work-life balance.
- Equity: Grants are available for eligible roles, subject to board approval.
Company info
- Pilots don't train with real passengers.
- Actors don't rehearse with real audiences.
- Yet, the most consequential decisions in society are often pushed straight to production.
- Simile is changing that.
- We have built the first AI simulation of society, populated by generative agents based on real humans.
- Our research pioneered the field of AI-based simulation, proving it is possible to model human behavior with high accuracy.
- Today, we are developing a Foundation Model to predict human behavior in any situation, at any scale.
- We are backed by $100M in funding led by Index Ventures, with participation from Hanabi, A*, Bain Capital Ventures, and AI visionaries including Andrej Karpathy, Fei-Fei Li, Adam D'Angelo, and Guillermo Rauch.
- We are hiring across several forms of expertise.
- Help the company reason about sampling error, uncertainty, calibration, margin of error, representativeness, and what "ground truth" means when human behavior is inherently noisy.
- Evaluate new model versions, diagnose regressions, identify priority areas for model-improvement cycles, and maintain stable eval suites that represent capabilities customers actually care about.
- We are hiring a team with complementary strengths.
Equal opportunity
- Simile is an equal opportunity workplace.
- We welcome applicants of all backgrounds and identities, valuing an environment where everyone can contribute authentically.
- Accommodations: If you require support or reasonable accommodations during the application process due to a disability, please let us know.
- We are happy to assist.
This listing is sourced directly from Simile's careers page and normalized into a canonical job model.