Anthropic

Anthropic

Staff+ Software Engineer, Safeguards ML Infrastructure

San Francisco, CA · Staff+

H1B sponsorship available$320k-$485kDetected 7 days ago
PythonRustDistributed SystemsBackend DevelopmentAWSGCPCloud PlatformsMachine LearningNLPLLMsIncident ResponseLogisticsResearchCommunication

About the role

  • We're growing the team and looking for Software Engineers with deep experience owning production infrastructure at scale.
  • The ideal candidate has built and operated large-scale distributed systems under real production pressure and built platforms, tooling, and infrastructure that other engineers depend on.
  • What we prioritize is distributed backend systems expertise and a track record of ownership over the production environment.

Responsibilities

  • Design, build, and deploy backend services that are critical safety pieces on the token sampling and generation path.
  • Own and operate the production serving infrastructure for those services across multiple deployment platforms (1P, AWS Bedrock, GCP Vertex).
  • Define and maintain SLOs, build observability and alerting systems, and lead incident response for infrastructure on the critical path of every Claude request
  • Reduce oncall and onduty toil by building automation, tooling, and self-serve workflows that minimize manual operations. Be the first user of the systems you build, running them for real workloads yourself before other teams depend on them.
  • Build and maintain a safety registry with full provenance -- tracking what is running in production, on which model, and when and by whom it was deployed.
  • Implement automated post-deploy validation to ensure correctness is consistent across platforms.
  • Have designed, built, and operated high QPS systems at global scale.
  • Experience building deployment and rollout systems with canary analysis, automated validation, or progressive rollout controls.

Requirements

  • Years of experience required will correlate with the internal job level requirements for the position
  • Familiarity with ML research or transformer architectures is not required -- you will learn that on the job.
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Nice to have

  • Are proficient in Python
  • experience with Rust is a plus but not required.

Skills

  • Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements.
  • Have hands-on experience deploying and operating on cloud platforms (AWS, GCP) at scale

Compensation

  • $320,000 - $485,000 USD

Benefits

  • Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience

Visa & Work Authorization

  • However, we aren't able to successfully sponsor visas for every role and every candidate.
  • But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
  • We do sponsor visas!

This listing is sourced directly from Anthropic's careers page and normalized into a canonical job model.