PMs for Hire

PMs for Hire

Research Engineer, Safeguards Labs

San Francisco, CA | New York City, NY

H1B sponsorship available$350k-$850kDetected 29 days ago
PythonMachine LearningNLPLLMsCybersecurityDetection EngineeringLogisticsResearchCommunication

About the role

  • We're hiring research engineers to define and execute the Labs research agenda.
  • The team is small and being built deliberately around a roughly 3:1 mix of researchers to software engineers, so each person has substantial latitude over what they work on and high leverage on the team's direction.
  • Contribute to a broader research portfolio investigating methods for detecting abusive behavior in chat-based or agentive workflows, and for training the model to robustly refrain from dangerous responses or behaviors without over-refusing.

Responsibilities

  • Lead and contribute to research projects investigating new methods for detecting misuse of Claude, identifying malicious organizations and accounts, strengthening model safeguards, and other safety needs.
  • Design and run offline analyses over model usage data to surface abuse patterns, build classifiers and detection systems, and evaluate their effectiveness.
  • Develop and iterate on prototypes that could eventually feed signals into the real-time safeguards path, partnering with engineers on tech transfer.
  • Build evaluations and methodologies for measuring whether safeguards actually work, including in agentic settings.
  • Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
  • You'll scope your own projects, run experiments end-to-end, and decide when an idea is ready to hand off to a production team - or when to kill it and move on.

Requirements

  • Have a track record of independently driving research projects from ambiguous problem statements to concrete results, ideally in AI, ML, security, integrity, or a related technical field.
  • Years of experience required will correlate with the internal job level requirements for the position

Nice to have

  • Knowledge of evaluation methodologies for language models and experience designing evals.
  • Experience with agentic environments and evaluating model behavior in them.
  • Background in trust and safety, integrity, fraud detection, threat intelligence, or adversarial ML.
  • Experience with red teaming, jailbreak research, or interpretability methods like steering vectors.
  • A history of taking research prototypes and transferring them into production systems.
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Compensation

  • $350,000 - $850,000 USD
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
  • However, some roles may require more time in our offices.

Benefits

  • Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience

Company info

  • Safeguards Labs is a new team operating at the intersection of research and engineering, chartered to investigate novel safety methods that protect Claude and the people who use it.
  • We prototype new approaches to safe models, usage safeguards, and production safety - pressure-testing ideas through offline analysis and subsets of traffic before they graduate into production systems run by our partner Safeguards teams.
  • Our work overlaps closely with account abuse, model behavior safeguards, and other safeguard subteams, and we serve as a research arm that can take on ambitious, ambiguous problems and turn them into deployed defenses.
  • Anthropic's mission is to create reliable, interpretable, and steerable AI systems.

Visa & Work Authorization

  • However, we aren't able to successfully sponsor visas for every role and every candidate.
  • But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
  • We do sponsor visas!

This listing is sourced directly from PMs for Hire's careers page and normalized into a canonical job model.