Inception

Inception

Member of Technical Staff, Model Evaluation

San Mateo, USA · Staff+

Sponsorship not specifiedDetected 134 days ago
PythonGitAWSGCPAzureCloud PlatformsDockerHelmCI/CDMachine LearningData EngineeringLLMsMLOpsStatisticsResearchExperimental DesignCommunication

About the role

  • We seek experienced engineers and scientists to develop the evaluation metrics and systems that drive frontier LLM performance.
  • You'll design the frameworks that tell us whether our models are improving and ensure they perform reliably at scale in production.

Responsibilities

  • Design, develop, and maintain robust evaluation frameworks and benchmarks for measuring LLM performance across diverse tasks and domains.
  • Define and implement quantitative metrics that capture model quality, safety, reliability, and regression detection.
  • Build scalable, automated evaluation pipelines that integrate into model training and deployment workflows.
  • Partner with product and customer-facing teams to translate real-world use cases into meaningful evaluation criteria.
  • Solid foundation in statistics, experimental design, and hypothesis testing.

Requirements

  • Experience with data engineering, large-scale data labeling, or synthetic data generation for evaluation purposes.

Nice to have

  • Conduct rigorous statistical analysis of model outputs to identify failure modes, biases, and performance gaps.
  • At least 2 years of experience in ML evaluation, applied ML research, or a related engineering role.
  • Proficiency in Python and ML frameworks such as PyTorch.
  • Experience designing and implementing evaluation metrics and benchmarks for generative models.
  • Experience with version control (Git) and containerization (Docker).
  • Excellent communication skills with the ability to distill complex evaluation results into actionable insights.
  • Preferred Skills Experience with human-in-the-loop evaluation systems (Likert-scale annotation, pairwise preference ranking, red-teaming).
  • Familiarity with LLM safety and alignment evaluation (toxicity, hallucination detection, factual grounding).

Benefits

  • Qualifications BS/MS/PhD in Computer Science, Machine Learning, Statistics, or a related field (or equivalent experience).
  • Strong understanding of LLM fundamentals (autoregressive generation, instruction tuning, RLHF, in-context learning, decoding strategies).

This listing is sourced directly from Inception's careers page and normalized into a canonical job model.