Baseten

Baseten

Manager, Cloud Platform & Site Reliability

San Francisco · Senior

Sponsorship not specifiedDetected 66 days ago
Node.jsDistributed SystemsGitCloud PlatformsKubernetesTerraformHelmCI/CDGitHub ActionsJenkinsPrometheusGrafanaSite Reliability EngineeringPlatform EngineeringMachine LearningSparkIncident ResponseCustomer SuccessResearchCommunicationCollaborationMentoring

About the role

  • This role requires someone who can zoom out to set org-level direction while remaining technically credible enough to engage meaningfully in architectural decisions across Kubernetes, multi-cloud infrastructure, and reliability engineering.

Responsibilities

  • Lead, grow, and develop team leads across the Cloud Platform and Site Reliability Engineering orgs, building a culture of ownership, technical excellence, and continuous improvement.
  • Own the reliability posture of the platform end-to-end, establishing and enforcing org-wide standards for SLOs/SLIs, incident response, observability-as-code, runbooks, and post-incident reviews.
  • Drive cross-functional collaboration with product, engineering, and customer-facing teams to ensure infrastructure capabilities and reliability investments align with product goals and enterprise customer requirements.
  • Partner with forward-deployed and customer success teams to support enterprise accounts with strict SLAs and complex infrastructure requirements.
  • Demonstrate accountability, pride of ownership, and high standards - and expect the same from your leads and their teams.
  • We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital.
  • Join us and help build the platform engineers turn to to ship AI products.

Requirements

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
  • Proven experience managing managers and leading multiple high-performing infrastructure, platform, or SRE teams in a fast-paced, high-growth environment.
  • Deep technical expertise in Kubernetes (multi-cloud across EKS, GKE, or similar), cloud infrastructure, and distributed systems, with the ability to engage credibly in architectural and operational decisions.
  • familiarity with GitOps workflows (e.g., Flux CD, ArgoCD, Helm).
  • Experience owning incident management and enterprise SLAs at scale, including executive-level communication during high-severity incidents and rigorous post-incident follow-through.
  • Familiarity with running high-performance AI models and workloads, including troubleshooting ML pipelines from preprocessing through inference and serving.
  • Experience with GPU infrastructure, including fractional GPU provisioning and multi-node model serving (e.g., on H100s or B200s).
  • Experience with incident management platforms (e.g., incident.io, PagerDuty) and building AI-assisted tooling for incident triage and response.

Compensation

  • Competitive compensation, including meaningful equity.

Benefits

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
  • If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
  • By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.

Company info

  • At Baseten, we are committed to fostering a diverse and inclusive workplace.

This listing is sourced directly from Baseten's careers page and normalized into a canonical job model.