Baseten

Baseten

Site Reliability Engineer

San Francisco

Sponsorship not specifiedDetected 72 days ago
KubernetesTerraformHelmPrometheusGrafanaSite Reliability EngineeringMachine LearningSparkIncident ResponseResearchArgo CD

About the role

  • As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform.
  • You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company.
  • You'll work on projects like these as part of the SRE team:

Responsibilities

  • Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking.
  • Build and maintain observability infrastructure - metrics, logging, dashboards, and alerting - as code.
  • Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution.
  • Navigate ambiguity, make principled tradeoffs, and avoid unnecessary complexity in the systems you build and the processes you define.
  • We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital.
  • Join us and help build the platform engineers turn to to ship AI products.
  • You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale - and that empower the broader organization to operate confidently.

Nice to have

  • Experience with infrastructure-as-code (Terraform, Helm) and GitOps workflows (Flux CD, ArgoCD).
  • Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis.
  • Comfort working at the intersection of engineering and operations - you write code, but you also think deeply about process, escalation paths, and operational leverage.
  • Familiarity with incident management platforms (incident.io http://incident.io or similar) is a plus.
  • No prior ML experience required, but curiosity about how ML models are deployed and served at scale will serve you well.
  • Extensive hands-on experience with Kubernetes (multi-cloud experience across EKS, GKE, or similar is a strong plus).
  • Strong foundation in observability tooling: metrics (VictoriaMetrics, Prometheus), logging (Loki, ELK), dashboards (Grafana), and alerting pipelines.
  • Observability-as-code experience is a plus.

Compensation

  • Competitive compensation, including meaningful equity.

Benefits

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
  • If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
  • By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.

Company info

  • learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company.
  • At Baseten, we are committed to fostering a diverse and inclusive workplace.

This listing is sourced directly from Baseten's careers page and normalized into a canonical job model.