Sanas
Research Scientist (Model Evaluation)
Palo Alto, CA
Sponsorship not specifiedDetected 16 days ago
PythonMachine LearningPyTorchNLPRoadmappingCustomer SuccessSystems EngineeringResearchCommunication
About the role
- Progress in speech AI is only as meaningful as our ability to measure it.
- At Sanas, model quality spans dimensions that automated metrics struggle to capture - accent naturalness, perceptual clarity, speaker identity preservation, noise suppression without speech distortion, translation fluency under real-world disfluency.
- This role sits at the intersection of research, product, and infrastructure - and directly shapes how every model team at Sanas measures progress.
Responsibilities
- Build evaluation systems that bridge automated metrics and human judgment - designing listening studies, MOS/MUSHRA protocols, and preference tests that are statistically rigorous and operationally scalable.
- Evaluation infrastructure & tooling Build and maintain automated evaluation pipelines that run continuously against model checkpoints - surfacing regressions early and tracking quality trends across training runs.
- Develop reference-based and reference-free metrics calibrated to
- Build tooling that allows research scientists and ML engineers to run rigorous ablations, compare model versions, and understand quality tradeoffs without needing to design the evaluation from scratch each time.
- Human evaluation & research Design and operate human evaluation programs - listener panels, crowdsourced annotation, and expert evaluator workflows - that produce reliable signal on dimensions automated metrics cannot capture.
- Develop novel quantitative metrics for subjective and perceptual qualities: accent similarity, naturalness, speaker identity preservation, intelligibility under noise, and translation fluency in spoken-language domains.
- Develop reference-based and reference-free metrics calibrated to Sanas's specific model tasks: SI-SDR, PESQ, STOI, DNSMOS, speaker similarity, WER delta, COMET, and task-specific custom metrics where off-the-shelf measures fall short.
Requirements
- About the Role Progress in speech AI is only as meaningful as our ability to measure it.
Skills
- Proficiency in Python and PyTorch or equivalent.
- Ability to take open-ended research questions and translate them into concrete, measurable evaluation systems that run reliably at scale.
- Familiarity with real-time or streaming model evaluation - latency-quality tradeoffs, codec-degraded audio, telephony channel conditions.
- Background in psychoacoustics or perceptual audio quality - understanding of how humans perceive speech naturalness, noise, and distortion.
- Published research at INTERSPEECH, ICASSP, ACL, EMNLP, or equivalent venues on evaluation methodology, speech quality, or related topics.
Benefits
- Bonus Experience evaluating models across multiple speech tasks - ASR, TTS, speech enhancement, speaker verification, or machine translation.
Company info
- Cross-functional impact Work closely with ML research, product, and customer success teams to ensure evaluation reflects what customers actually experience - not just what lab conditions optimize for.
This listing is sourced directly from Sanas's careers page and normalized into a canonical job model.