Anthropic

Anthropic

Staff + Sr. Software Engineer, Cloud Inference Launch Engineering

San Francisco, CA · Staff+

H1B sponsorship available$320k-$485kDetected 7 days ago
PythonRustDistributed SystemsAWSGCPAzureCloud PlatformsKubernetesCI/CDMachine LearningNLPLLMsLogisticsLoad BalancingResearchCommunicationCollaboration

About the role

  • We're responsible for every inference change - model launches, performance improvements, safeguard integrations - landing on cloud platforms with correctness, performance, and reliability intact.
  • This directly determines how fast frontier models and features ship to every cloud platform, and how quickly performance wins reach production - reclaiming capacity at a time when compute is our scarcest resource.

Responsibilities

  • Identify and dive deep on the gaps that make inference behave differently across first-party and CSPs - config drift, observability, deployment patterns, hard cross-platform bugs - and fix them at the source rather than building platform-specific workarounds
  • Design, build, and own the CI/CD infrastructure for the inference server and load balancer across cloud platforms, with shadow traffic, performance baselines (throughput and latency), and correctness checks that catch regressions before production
  • Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation based on real-world production workloads
  • Have a track record of building automation or test infrastructure that measurably improved release velocity or reliability
  • Have experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration
  • Thrive in cross-functional collaboration with both internal teams and external partners
  • Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

Requirements

  • prior inference or ML experience is not required
  • Have significant software engineering experience, with a strong background in high-performance, large-scale distributed systems serving millions of users
  • Years of experience required will correlate with the internal job level requirements for the position
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Nice to have

  • LLM inference optimization, batching, and caching strategies
  • Capacity-constrained scheduling or shared-resource test infrastructure
  • Solid understanding of multi-region deployments, request routing, load balancing, global traffic management
  • Proficiency in Python or Rust

Compensation

  • $320,000 - $485,000 USD

Benefits

  • Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience

Company info

  • Drive down merge-to-production cycle time by making validation faster, more parallel, and cost-effective enough to run on the same constrained accelerator pool that serves customers, without trading away reliability
  • Anthropic's mission is to create reliable, interpretable, and steerable AI systems.

Visa & Work Authorization

  • However, we aren't able to successfully sponsor visas for every role and every candidate.
  • But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
  • We do sponsor visas!

This listing is sourced directly from Anthropic's careers page and normalized into a canonical job model.