Baseten
Manager, Cloud Platform & Site Reliability
San Francisco · Senior
Sponsorship not specifiedDetected 66 days ago
Node.jsDistributed SystemsGitCloud PlatformsKubernetesTerraformHelmCI/CDGitHub ActionsJenkinsPrometheusGrafanaSite Reliability EngineeringPlatform EngineeringMachine LearningSparkIncident ResponseCustomer SuccessResearchCommunicationCollaborationMentoring
About the role
- This role requires someone who can zoom out to set org-level direction while remaining technically credible enough to engage meaningfully in architectural decisions across Kubernetes, multi-cloud infrastructure, and reliability engineering.
Responsibilities
- Lead, grow, and develop team leads across the Cloud Platform and Site Reliability Engineering orgs, building a culture of ownership, technical excellence, and continuous improvement.
- Own the reliability posture of the platform end-to-end, establishing and enforcing org-wide standards for SLOs/SLIs, incident response, observability-as-code, runbooks, and post-incident reviews.
- Drive cross-functional collaboration with product, engineering, and customer-facing teams to ensure infrastructure capabilities and reliability investments align with product goals and enterprise customer requirements.
- Partner with forward-deployed and customer success teams to support enterprise accounts with strict SLAs and complex infrastructure requirements.
- Demonstrate accountability, pride of ownership, and high standards - and expect the same from your leads and their teams.
- We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital.
- Join us and help build the platform engineers turn to to ship AI products.
Requirements
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
- Proven experience managing managers and leading multiple high-performing infrastructure, platform, or SRE teams in a fast-paced, high-growth environment.
- Deep technical expertise in Kubernetes (multi-cloud across EKS, GKE, or similar), cloud infrastructure, and distributed systems, with the ability to engage credibly in architectural and operational decisions.
- familiarity with GitOps workflows (e.g., Flux CD, ArgoCD, Helm).
- Experience owning incident management and enterprise SLAs at scale, including executive-level communication during high-severity incidents and rigorous post-incident follow-through.
- Familiarity with running high-performance AI models and workloads, including troubleshooting ML pipelines from preprocessing through inference and serving.
- Experience with GPU infrastructure, including fractional GPU provisioning and multi-node model serving (e.g., on H100s or B200s).
- Experience with incident management platforms (e.g., incident.io, PagerDuty) and building AI-assisted tooling for incident triage and response.
Compensation
- Competitive compensation, including meaningful equity.
Benefits
- Competitive compensation, including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
- If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
- By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.
Company info
- At Baseten, we are committed to fostering a diverse and inclusive workplace.
Apply directly at Baseten →Create a free account for alerts like thisView Baseten immigration profile
This listing is sourced directly from Baseten's careers page and normalized into a canonical job model.