Simile

Simile

Evaluations - Member of Technical Staff

San Francisco · Staff+

Sponsorship not specified$200k-$400kDetected 21 days ago
PythonSQLMachine LearningData ScienceLLMsStatisticsA/B TestingResearchExperimental Design

About the role

  • You will help shape what Simile measures, the quality bars we defend, and how evaluation evidence guides model, product, and customer decisions.
  • Evaluation at Simile brings together model evals, statistics, behavioral science, research methodology, product quality, and human judgment.
  • Our models simulate people, populations, markets, and groups, which means our evals must reason about distributions, noisy human ground truth, uncertainty, qualitative outputs, behavioral data, and customer decision-making.

Responsibilities

  • Design evals, metrics, rubrics, datasets, dashboards, and workflows that measure whether Simile's models are accurately predicting human behavior across customer use cases, populations, question types, and decision contexts.
  • Build evals for qualitative responses, retrieval, survey generation, AI-generated research reports, customer-facing outputs, and other product surfaces where model quality directly shapes customer trust.
  • Develop rigorous ways to compare simulated responses against human data, customer studies, Simile-collected ground truth, and behavioral datasets.
  • Experience building model evaluation dashboards, regression suites, release gates, benchmark sets, model comparison workflows, or systems that help ML teams decide where to focus and when to ship.
  • Our hiring journey is designed to help both sides align on fit, working style, and expectations.

Requirements

  • Evaluation Taste: You have strong intuition for what makes an eval meaningful, robust, and decision-relevant.
  • You can explain what an eval measures, what it does not measure, how it can be gamed, and why it should or should not affect a model or product decision.
  • You do not need to be a modeling specialist, but you can read model outputs, understand modeling team needs, and reason about whether a model change actually improved the thing we care about.
  • You are comfortable working with data and automation tools such as Python, SQL, R, notebooks, LLM APIs, and agentic coding tools such as Codex, Claude Code, Cursor, or equivalent systems.
  • You know how to move quickly while still validating outputs, catching errors, and planning for the long-term..
  • Survey Methodology and Statistics: Experience with sampling, weighting, margin of error, power analysis, uncertainty quantification, Bayesian modeling, causal inference, psychometrics, polling, or measurement theory.
  • Multi-Agent or Group Behavior: Interest or experience in modeling group conversation, deliberation, focus groups, juries, committees, polarization, collective decision-making, or social influence.

Skills

  • Final offers are based on experience, specialized skills, interview performance, and relevant training.

Compensation

  • $200,000 - $400,000 USD
  • At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.

Benefits

  • At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.
  • Comprehensive medical, dental, and vision coverage.
  • Flexible time off policies to support work-life balance.
  • Equity: Grants are available for eligible roles, subject to board approval.

Company info

  • Pilots don't train with real passengers.
  • Actors don't rehearse with real audiences.
  • Yet, the most consequential decisions in society are often pushed straight to production.
  • Simile is changing that.
  • We have built the first AI simulation of society, populated by generative agents based on real humans.
  • Our research pioneered the field of AI-based simulation, proving it is possible to model human behavior with high accuracy.
  • Today, we are developing a Foundation Model to predict human behavior in any situation, at any scale.
  • We are backed by $100M in funding led by Index Ventures, with participation from Hanabi, A*, Bain Capital Ventures, and AI visionaries including Andrej Karpathy, Fei-Fei Li, Adam D'Angelo, and Guillermo Rauch.
  • We are hiring across several forms of expertise.
  • Help the company reason about sampling error, uncertainty, calibration, margin of error, representativeness, and what "ground truth" means when human behavior is inherently noisy.
  • Evaluate new model versions, diagnose regressions, identify priority areas for model-improvement cycles, and maintain stable eval suites that represent capabilities customers actually care about.
  • We are hiring a team with complementary strengths.

Equal opportunity

  • Simile is an equal opportunity workplace.
  • We welcome applicants of all backgrounds and identities, valuing an environment where everyone can contribute authentically.
  • Accommodations: If you require support or reasonable accommodations during the application process due to a disability, please let us know.
  • We are happy to assist.

This listing is sourced directly from Simile's careers page and normalized into a canonical job model.