Baseten
Site Reliability Engineer
San Francisco
Sponsorship not specifiedDetected 72 days ago
KubernetesTerraformHelmPrometheusGrafanaSite Reliability EngineeringMachine LearningSparkIncident ResponseResearchArgo CD
About the role
- As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform.
- You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company.
- You'll work on projects like these as part of the SRE team:
Responsibilities
- Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking.
- Build and maintain observability infrastructure - metrics, logging, dashboards, and alerting - as code.
- Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution.
- Navigate ambiguity, make principled tradeoffs, and avoid unnecessary complexity in the systems you build and the processes you define.
- We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital.
- Join us and help build the platform engineers turn to to ship AI products.
- You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale - and that empower the broader organization to operate confidently.
Nice to have
- Experience with infrastructure-as-code (Terraform, Helm) and GitOps workflows (Flux CD, ArgoCD).
- Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis.
- Comfort working at the intersection of engineering and operations - you write code, but you also think deeply about process, escalation paths, and operational leverage.
- Familiarity with incident management platforms (incident.io http://incident.io or similar) is a plus.
- No prior ML experience required, but curiosity about how ML models are deployed and served at scale will serve you well.
- Extensive hands-on experience with Kubernetes (multi-cloud experience across EKS, GKE, or similar is a strong plus).
- Strong foundation in observability tooling: metrics (VictoriaMetrics, Prometheus), logging (Loki, ELK), dashboards (Grafana), and alerting pipelines.
- Observability-as-code experience is a plus.
Compensation
- Competitive compensation, including meaningful equity.
Benefits
- Competitive compensation, including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
- If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
- By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.
Company info
- learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company.
- At Baseten, we are committed to fostering a diverse and inclusive workplace.
Apply directly at Baseten →Create a free account for alerts like thisView Baseten immigration profile
This listing is sourced directly from Baseten's careers page and normalized into a canonical job model.