Anthropic
Research Engineer, Model Evaluations
Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY
H1B sponsorship available$500k-$850kDetected 7 days ago
PythonDistributed SystemsMachine LearningData EngineeringData VisualizationLLMsStatisticsLogisticsResearchExperimental DesignLeadershipCommunication
About the role
- Your work will turn ambiguous notions of "intelligence" into clear, defensible metrics that researchers, leadership, and the public can rely on.
- The goal is to make Anthropic the leader in extremely well-characterized AI systems, with performance that is exhaustively measured and validated across the tasks that matter.
- Debug anomalous eval results mid-training-run, determine whether the cause is a model change or an infrastructure issue, and communicate the answer clearly under time pressure
Responsibilities
- Design and run new evaluations of Claude's capabilities - reasoning, agentic behavior, knowledge, safety - and produce visualizations that make the results legible to researchers and decision-makers
- Build and harden the distributed eval execution platform so hundreds of evals run reliably against checkpoints throughout production RL training runs
- Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations
- Partner with research teams across the full lifecycle of a new capability - from defining what to measure to interpreting results as training progresses
- Experience building or operating distributed systems, data pipelines, or other infrastructure that needs to be reliable at scale
- Comfort operating in an on-call or production-support capacity when training runs are live
- Partner with a research team on a new capability area, helping them articulate what "good" looks like and translating that into measurable artifacts
Requirements
- Years of experience required will correlate with the internal job level requirements for the position
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
Nice to have
- Hands-on experience using large language models such as Claude, including prompting, sampling, and scaffolding
- Experience with observability, monitoring, or experiment-tracking systems
- Experience with large-scale dataset sourcing, curation, and processing
- Experience running or supporting ML training infrastructure
- A bias toward picking up slack and operating flexibly across team boundaries
- Enjoy pair programming - we love to pair
- Representative projects
- Diagnose a mid-training regression: an eval suite returns anomalous numbers, and you need to determine within hours whether it's the model, the harness, the data, or the infrastructure
Compensation
- $500,000 - $850,000 USD
Benefits
- Own the dashboards researchers and leadership use to monitor model health during training, improving signal-to-noise, reducing latency, and making regressions impossible to miss
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience
Visa & Work Authorization
- However, we aren't able to successfully sponsor visas for every role and every candidate.
- But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
- We do sponsor visas!
Apply directly at Anthropic →Create a free account for alerts like thisView Anthropic immigration profile
This listing is sourced directly from Anthropic's careers page and normalized into a canonical job model.