Perplexity AI

Perplexity AI

Member of Technical Staff (Data Scientist, Evals)

San Francisco · Staff+

Sponsorship not specifiedDetected 23 days ago
PythonSQLDatabricksAWSMachine LearningData ScienceLLMsResearchLeadership

About the role

  • Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and our specialized data sources.
  • We aim to use the latest models as they are released, but the intelligence frontier is a jagged one, and popular benchmarks do not effectively cover our use cases.

Responsibilities

  • Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness
  • Design evaluation sets and methods specifically to measure the impact of tool calls (particularly web search retrieval) on the final answer's quality
  • Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices
  • In this role, you will build specialized evals to improve answer quality across Perplexity, covering search-based LLM answers and other scenarios popular with our users.

Requirements

  • Strong proficiency in Python and SQL (expected to write production-grade code)
  • PhD or MS in a technical field or equivalent experience
  • 4+ years of experience in data science or machine learning
  • Experience building within a modern cloud data stack, specifically AWS and Databricks
  • Comfortable with agentic coding workflows and using AI-assisted development tools to iterate faster

Nice to have

  • 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups
  • Prior experience working on customer-facing web products or consumer apps, with real user traffic at scale
  • A strong research background, with experience applying research methods to real-world ML problems

This listing is sourced directly from Perplexity AI's careers page and normalized into a canonical job model.