Upstart

Upstart

Senior Engineering Manager, Site Reliability

United States | Remote · Senior

Sponsorship not specified$195k-$270kDetected 4 days ago
Distributed SystemsAWSCloud PlatformsKubernetesPrometheusGrafanaDatadogSite Reliability EngineeringPlatform EngineeringSparkA/B TestingIncident ResponseAccessibilityCadenceLeadershipCommunication

About the role

  • Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function
  • Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures
  • Set a high bar for technical quality, operating rigor, and executive communication

Responsibilities

  • Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering
  • Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership
  • Develop engineers and leaders who can independently own complex reliability initiatives
  • Identify recurring failure patterns and drive systemic solutions across teams
  • Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction

Requirements

  • 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering
  • Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems
  • Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes
  • Experience leading high severity incident response and improving incident management practices at scale

Nice to have

  • Experience operating large scale, highly available distributed systems
  • Experience implementing or evolving service-level objectives and error-budget practices
  • Experience with observability platforms such as Datadog, Grafana, Prometheus, OpenTelemetry, or similar technologies
  • Familiarity with Kubernetes, AWS, and modern cloud native architectures
  • Experience supporting major platform or architectural transitions
  • Experience establishing executive level reliability reporting and operating reviews

Skills

  • Individual pay is also determined by job-related skills, experience, and relevant education or training.

Compensation

  • At Upstart, your base pay is one part of your total compensation package.
  • The anticipated base salary for this position is expected to be within the below range.
  • United States | Remote - Anticipated Base Salary Range
  • $195,300 - $270,400 USD

Benefits

  • In addition, Upstart provides employees with target bonuses, equity compensation, and generous benefits packages (including medical, dental, vision, and 401k).
  • At Upstart, our benefits are designed to support your health, financial well-being, family, and personal growth.
  • Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts
  • Incident Management and Learning

Company info

  • The Site Reliability Engineering (SRE) team enables Upstart's engineering organization to operate reliable, observable, and resilient systems at scale.
  • The team owns company-wide incident management practices, reliability standards, operational readiness, and the capabilities that help engineering teams identify, respond to, and learn from production issues.
  • Our goal is to make reliability an integrated part of how software is designed, delivered, and operated.
  • We are building a model where engineering teams have the trusted signals, automated safeguards, and operational practices needed to move quickly while protecting our customers and business.
  • The team advances observability, incident detection and response, service level objectives, operational readiness, and systemic improvements based on incident learnings.
  • SRE partners across product engineering, infrastructure, security, and platform teams to improve reliability at scale.
  • As a digital first company, the majority of your work can be accomplished remotely.
  • At Upstart, your base pay is one part of your total compensation package.

This listing is sourced directly from Upstart's careers page and normalized into a canonical job model.