Anthropic

Anthropic

Research Engineer, Model Evaluations

Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY

H1B sponsorship available$500k-$850kDetected 7 days ago
PythonDistributed SystemsMachine LearningData EngineeringData VisualizationLLMsStatisticsLogisticsResearchExperimental DesignLeadershipCommunication

About the role

  • Your work will turn ambiguous notions of "intelligence" into clear, defensible metrics that researchers, leadership, and the public can rely on.
  • The goal is to make Anthropic the leader in extremely well-characterized AI systems, with performance that is exhaustively measured and validated across the tasks that matter.
  • Debug anomalous eval results mid-training-run, determine whether the cause is a model change or an infrastructure issue, and communicate the answer clearly under time pressure

Responsibilities

  • Design and run new evaluations of Claude's capabilities - reasoning, agentic behavior, knowledge, safety - and produce visualizations that make the results legible to researchers and decision-makers
  • Build and harden the distributed eval execution platform so hundreds of evals run reliably against checkpoints throughout production RL training runs
  • Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations
  • Partner with research teams across the full lifecycle of a new capability - from defining what to measure to interpreting results as training progresses
  • Experience building or operating distributed systems, data pipelines, or other infrastructure that needs to be reliable at scale
  • Comfort operating in an on-call or production-support capacity when training runs are live
  • Partner with a research team on a new capability area, helping them articulate what "good" looks like and translating that into measurable artifacts

Requirements

  • Years of experience required will correlate with the internal job level requirements for the position
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Nice to have

  • Hands-on experience using large language models such as Claude, including prompting, sampling, and scaffolding
  • Experience with observability, monitoring, or experiment-tracking systems
  • Experience with large-scale dataset sourcing, curation, and processing
  • Experience running or supporting ML training infrastructure
  • A bias toward picking up slack and operating flexibly across team boundaries
  • Enjoy pair programming - we love to pair
  • Representative projects
  • Diagnose a mid-training regression: an eval suite returns anomalous numbers, and you need to determine within hours whether it's the model, the harness, the data, or the infrastructure

Compensation

  • $500,000 - $850,000 USD

Benefits

  • Own the dashboards researchers and leadership use to monitor model health during training, improving signal-to-noise, reducing latency, and making regressions impossible to miss
  • Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience

Visa & Work Authorization

  • However, we aren't able to successfully sponsor visas for every role and every candidate.
  • But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
  • We do sponsor visas!

This listing is sourced directly from Anthropic's careers page and normalized into a canonical job model.