PMs for Hire
Research Engineer, Safeguards Labs
San Francisco, CA | New York City, NY
H1B sponsorship available$350k-$850kDetected 29 days ago
PythonMachine LearningNLPLLMsCybersecurityDetection EngineeringLogisticsResearchCommunication
About the role
- We're hiring research engineers to define and execute the Labs research agenda.
- The team is small and being built deliberately around a roughly 3:1 mix of researchers to software engineers, so each person has substantial latitude over what they work on and high leverage on the team's direction.
- Contribute to a broader research portfolio investigating methods for detecting abusive behavior in chat-based or agentive workflows, and for training the model to robustly refrain from dangerous responses or behaviors without over-refusing.
Responsibilities
- Lead and contribute to research projects investigating new methods for detecting misuse of Claude, identifying malicious organizations and accounts, strengthening model safeguards, and other safety needs.
- Design and run offline analyses over model usage data to surface abuse patterns, build classifiers and detection systems, and evaluate their effectiveness.
- Develop and iterate on prototypes that could eventually feed signals into the real-time safeguards path, partnering with engineers on tech transfer.
- Build evaluations and methodologies for measuring whether safeguards actually work, including in agentic settings.
- Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
- You'll scope your own projects, run experiments end-to-end, and decide when an idea is ready to hand off to a production team - or when to kill it and move on.
Requirements
- Have a track record of independently driving research projects from ambiguous problem statements to concrete results, ideally in AI, ML, security, integrity, or a related technical field.
- Years of experience required will correlate with the internal job level requirements for the position
Nice to have
- Knowledge of evaluation methodologies for language models and experience designing evals.
- Experience with agentic environments and evaluating model behavior in them.
- Background in trust and safety, integrity, fraud detection, threat intelligence, or adversarial ML.
- Experience with red teaming, jailbreak research, or interpretability methods like steering vectors.
- A history of taking research prototypes and transferring them into production systems.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
Compensation
- $350,000 - $850,000 USD
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
- However, some roles may require more time in our offices.
Benefits
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience
Company info
- Safeguards Labs is a new team operating at the intersection of research and engineering, chartered to investigate novel safety methods that protect Claude and the people who use it.
- We prototype new approaches to safe models, usage safeguards, and production safety - pressure-testing ideas through offline analysis and subsets of traffic before they graduate into production systems run by our partner Safeguards teams.
- Our work overlaps closely with account abuse, model behavior safeguards, and other safeguard subteams, and we serve as a research arm that can take on ambitious, ambiguous problems and turn them into deployed defenses.
- Anthropic's mission is to create reliable, interpretable, and steerable AI systems.
Visa & Work Authorization
- However, we aren't able to successfully sponsor visas for every role and every candidate.
- But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
- We do sponsor visas!
Apply directly at PMs for Hire →Create a free account for alerts like thisView PMs for Hire immigration profile
This listing is sourced directly from PMs for Hire's careers page and normalized into a canonical job model.