Anthropic
Staff+ Software Engineer, Safeguards ML Infrastructure
San Francisco, CA · Staff+
H1B sponsorship available$320k-$485kDetected 7 days ago
PythonRustDistributed SystemsBackend DevelopmentAWSGCPCloud PlatformsMachine LearningNLPLLMsIncident ResponseLogisticsResearchCommunication
About the role
- We're growing the team and looking for Software Engineers with deep experience owning production infrastructure at scale.
- The ideal candidate has built and operated large-scale distributed systems under real production pressure and built platforms, tooling, and infrastructure that other engineers depend on.
- What we prioritize is distributed backend systems expertise and a track record of ownership over the production environment.
Responsibilities
- Design, build, and deploy backend services that are critical safety pieces on the token sampling and generation path.
- Own and operate the production serving infrastructure for those services across multiple deployment platforms (1P, AWS Bedrock, GCP Vertex).
- Define and maintain SLOs, build observability and alerting systems, and lead incident response for infrastructure on the critical path of every Claude request
- Reduce oncall and onduty toil by building automation, tooling, and self-serve workflows that minimize manual operations. Be the first user of the systems you build, running them for real workloads yourself before other teams depend on them.
- Build and maintain a safety registry with full provenance -- tracking what is running in production, on which model, and when and by whom it was deployed.
- Implement automated post-deploy validation to ensure correctness is consistent across platforms.
- Have designed, built, and operated high QPS systems at global scale.
- Experience building deployment and rollout systems with canary analysis, automated validation, or progressive rollout controls.
Requirements
- Years of experience required will correlate with the internal job level requirements for the position
- Familiarity with ML research or transformer architectures is not required -- you will learn that on the job.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
Nice to have
- Are proficient in Python
- experience with Rust is a plus but not required.
Skills
- Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements.
- Have hands-on experience deploying and operating on cloud platforms (AWS, GCP) at scale
Compensation
- $320,000 - $485,000 USD
Benefits
- Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience
Visa & Work Authorization
- However, we aren't able to successfully sponsor visas for every role and every candidate.
- But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
- We do sponsor visas!
Apply directly at Anthropic →Create a free account for alerts like thisView Anthropic immigration profile
This listing is sourced directly from Anthropic's careers page and normalized into a canonical job model.