Sayari

Sayari

Staff Applied Scientist - AI Evaluation & Trust

Remote - US · Staff+

Sponsorship not specified$195k-$225kDetected 41 days ago
Data EngineeringLLMsRAGStatisticsResearchExperimental Design

About the role

  • About Sayari: Sayari is the judgment infrastructure for trustworthy AI in economic security and commercial risk.
  • The Sayari Commercial World Model resolves 11.7B+ primary-source records from 250+ jurisdictions forming the ground truth of global commerce.
  • Trusted by U.S. Customs and Border Protection, HM Revenue & Customs, and Fortune 500 enterprises, Sayari is used by thousands of professionals across 35+ countries to secure supply chains and dismantle illicit networks.

Responsibilities

  • Lead the development of specialized "judge models," moving from general-purpose frontier models to architectures purpose-built for evaluation and failure mode detection.
  • Design and execute rigorous scoring pipelines and empirical threshold calibrations for agentic systems, including multi-turn conversation and Graph RAG reasoning.
  • Establish domain-specific evaluation frameworks that measure whether a system can perform the work of human experts rather than just passing general capability benchmarks.
  • Own the full lifecycle of evaluation data, from designing annotation infrastructure and protocols to deploying evaluation services into production.
  • Research and implement advanced techniques in Mixture-of-Experts (MoE) routing, expert specialization evaluation, and ensemble calibration.
  • Collaborate cross-functionally with Product, Data Engineering, and the SVP of AI to translate complex statistical uncertainty into clear, actionable product signals.
  • A Judgment Ontology, encoding over a decade of investigative tradecraft, and Superconductor, an agentic orchestration platform, deliver AI that reasons like an expert analyst, shows its work, and traces every finding to its source.
  • Mastery of statistics and experimental design, including significance testing, distribution analysis, and inter-rater reliability.

Requirements

  • 1-2+ years' experience focused on post-training activities
  • 1+ year experience creating benchmarks to evaluate LLMs
  • Experience with Mixture-of-Experts (MoE) systems, routing behavior, and expert specialization.

Compensation

  • $195,000 - $225,000 USD

Benefits

  • 100% fully paid medical, vision, and dental for employees and their dependents
  • Generous time off
  • we observe all US federal holidays, close our office for a winter break (12/24-12/31), in addition to granting 18 PTO days and 10 sick days
  • A strong commitment to diversity, equity, and inclusion
  • Eligibility to participate in additional benefits such as 401k match up to 5%, 100% paid life insurance (up to $100,000 coverage),, and parental leave
  • Limitless growth and learning opportunities
  • The target base salary for this position is $195,000-$225,000 plus company bonus and equity.
  • Final offer amounts are determined by multiple factors including location, local market variances, candidate experience and expertise, internal peer equity, and may vary from the amounts listed above.

Equal opportunity

  • equal opportunity employer and strongly encourages diverse candidates to apply.
  • We believe diversity and inclusion mean our team members should reflect the diversity of the United States.

This listing is sourced directly from Sayari's careers page and normalized into a canonical job model.