Sanas

Sanas

Research Scientist (Model Evaluation)

Palo Alto, CA

Sponsorship not specifiedDetected 26 days ago
PythonMachine LearningPyTorchNLPRoadmappingCustomer SuccessSystems EngineeringResearchCommunication

About the role

  • Progress in speech AI is only as meaningful as our ability to measure it.
  • At Sanas, model quality spans dimensions that automated metrics struggle to capture - accent naturalness, perceptual clarity, speaker identity preservation, noise suppression without speech distortion, translation fluency under real-world disfluency.
  • This role sits at the intersection of research, product, and infrastructure - and directly shapes how every model team at Sanas measures progress.

Responsibilities

  • Build evaluation systems that bridge automated metrics and human judgment - designing listening studies, MOS/MUSHRA protocols, and preference tests that are statistically rigorous and operationally scalable.
  • Evaluation infrastructure & tooling Build and maintain automated evaluation pipelines that run continuously against model checkpoints - surfacing regressions early and tracking quality trends across training runs.
  • Develop reference-based and reference-free metrics calibrated to
  • Build tooling that allows research scientists and ML engineers to run rigorous ablations, compare model versions, and understand quality tradeoffs without needing to design the evaluation from scratch each time.
  • Human evaluation & research Design and operate human evaluation programs - listener panels, crowdsourced annotation, and expert evaluator workflows - that produce reliable signal on dimensions automated metrics cannot capture.
  • Develop novel quantitative metrics for subjective and perceptual qualities: accent similarity, naturalness, speaker identity preservation, intelligibility under noise, and translation fluency in spoken-language domains.
  • Develop reference-based and reference-free metrics calibrated to Sanas's specific model tasks: SI-SDR, PESQ, STOI, DNSMOS, speaker similarity, WER delta, COMET, and task-specific custom metrics where off-the-shelf measures fall short.

Requirements

  • About the Role Progress in speech AI is only as meaningful as our ability to measure it.

Skills

  • Proficiency in Python and PyTorch or equivalent.
  • Ability to take open-ended research questions and translate them into concrete, measurable evaluation systems that run reliably at scale.
  • Familiarity with real-time or streaming model evaluation - latency-quality tradeoffs, codec-degraded audio, telephony channel conditions.
  • Background in psychoacoustics or perceptual audio quality - understanding of how humans perceive speech naturalness, noise, and distortion.
  • Published research at INTERSPEECH, ICASSP, ACL, EMNLP, or equivalent venues on evaluation methodology, speech quality, or related topics.

Benefits

  • Bonus Experience evaluating models across multiple speech tasks - ASR, TTS, speech enhancement, speaker verification, or machine translation.

Company info

  • Cross-functional impact Work closely with ML research, product, and customer success teams to ensure evaluation reflects what customers actually experience - not just what lab conditions optimize for.

This listing is sourced directly from Sanas's careers page and normalized into a canonical job model.