Simile

Simile

Evaluations Engineering - Member of Technical Staff

San Francisco · Staff+

Sponsorship not specified$200k-$400kDetected 5 days ago
Backend DevelopmentMachine LearningData EngineeringData ScienceLLMsStatisticsResearchCommunication

About the role

  • You will work across data and evaluation infrastructure, evaluation execution workflows, backend services, automation, and internal tooling.
  • Your initial focus will include streamlining how evaluations are run across models
  • strengthening evaluation versioning, data models, and access controls

Responsibilities

  • Develop the services, pipelines, and orchestration needed to run evaluations efficiently across datasets, model versions, populations, and use cases.
  • Design relational schemas, versioning, provenance, permissions, and quality controls that make evaluation results reproducible and trustworthy.
  • Partner with Evals and Data Operations to streamline customer validations, survey deployment, response ingestion, and the integration of new ground truth.
  • Build interfaces that help teams manage evals, compare models, investigate results, and identify regressions.
  • Our hiring journey is designed to help both sides align on fit, working style, and expectations.
  • As a Member of Technical Staff in Evaluations Engineering, you will build the systems that enable Simile to evaluate whether our simulations of human behavior are accurate, trustworthy, and improving over time.
  • Build evaluation execution infrastructure: Develop the services, pipelines, and orchestration needed to run evaluations efficiently across datasets, model versions, populations, and use cases.
  • Strengthen evaluation data systems: Design relational schemas, versioning, provenance, permissions, and quality controls that make evaluation results reproducible and trustworthy.
  • Automate validation and data collection: Partner with Evals and Data Operations to streamline customer validations, survey deployment, response ingestion, and the integration of new ground truth.
  • Develop evaluation tooling: Build interfaces that help teams manage evals, compare models, investigate results, and identify regressions.

Requirements

  • Strong Engineering Fundamentals: Several years of experience building and maintaining production-quality software, with sound judgment in system design, testing, debugging, and maintainability.
  • Ownership and Communication: A track record of independently driving important technical work and collaborating effectively across engineering, research, and operations.
  • Human Data Systems: Experience with labeling platforms, expert-review workflows, LLM-as-judge systems, grader calibration, or other human-in-the-loop evaluation methods.
  • Sensitive Data and Access Controls: Experience designing permissions, auditability, and data-governance systems for human or customer data.
  • Agentic Engineering: Experience using modern AI coding tools to accelerate development while independently testing and validating their output.
  • Several years of experience building and maintaining production-quality software, with sound judgment in system design, testing, debugging, and maintainability.
  • Experience with labeling platforms, expert-review workflows, LLM-as-judge systems, grader calibration, or other human-in-the-loop evaluation methods.
  • Experience using modern AI coding tools to accelerate development while independently testing and validating their output.

Skills

  • Final offers are based on experience, specialized skills, interview performance, and relevant training.

Compensation

  • $200,000 - $400,000 USD
  • At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.

Benefits

  • At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.
  • Comprehensive medical, dental, and vision coverage.
  • Flexible time off policies to support work-life balance.
  • Equity: Grants are available for eligible roles, subject to board approval.
  • application process due to a disability, please let us know.

Company info

  • Pilots don't train with real passengers.
  • Actors don't rehearse with real audiences.
  • Yet, the most consequential decisions in society are often pushed straight to production.
  • Simile is changing that.
  • We have built the first AI simulation of society, populated by generative agents based on real humans.
  • Our research pioneered the field of AI-based simulation, proving it is possible to model human behavior with high accuracy.
  • Today, we are developing a Foundation Model to predict human behavior in any situation, at any scale.
  • We are backed by $100M in funding led by Index Ventures, with participation from Hanabi, A*, Bain Capital Ventures, and AI visionaries including Andrej Karpathy, Fei-Fei Li, Adam D'Angelo, and Guillermo Rauch.
  • A track record of independently driving important technical work and collaborating effectively across engineering, research, and operations.
  • We do not expect one person to have all of these.
  • We are hiring a team with complementary strengths.
  • We are happy to assist.

Equal opportunity

  • Simile is an equal opportunity workplace.
  • We welcome applicants of all backgrounds and identities, valuing an environment where everyone can contribute authentically.
  • Accommodations: If you require support or reasonable accommodations during the application process due to a disability, please let us know.

This listing is sourced directly from Simile's careers page and normalized into a canonical job model.