Runpod

Runpod

Site Reliability Engineer

Remote, USA

Sponsorship not specified$150k-$200kDetected 60 days ago
PythonGoBashDistributed SystemsCI/CDLinuxPrometheusGrafanaSite Reliability EngineeringMachine LearningA/B TestingIncident ResponseLeadershipCommunicationProblem Solving

About the role

  • With over 500,000 developers worldwide and an annual recurring revenue run rate exceeding $120M, Runpod operates at the intersection of developer velocity and production-scale AI.
  • We value proactive problem solving, automation-first thinking, and strong ownership of production systems.
  • This role blends software engineering with production operations.

Responsibilities

  • Define and implement SLIs/SLOs for critical services
  • Lead incident response and coordinate cross-team mitigation efforts
  • Perform production readiness reviews for new services and features
  • Identify systemic risks and drive preventative improvements
  • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
  • Build internal tooling for reliability tracking and reporting
  • Build tools and scripts (Python, Go, Bash) to eliminate manual processes
  • Partner with engineering teams to improve system resilience
  • You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company.

Requirements

  • 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
  • Strong Linux systems and Networking expertise
  • Experience managing containerized production systems
  • Strong understanding of distributed systems and failure modes
  • Experience defining and managing SLIs/SLOs
  • Proven incident response and postmortem leadership experience
  • Experience with monitoring and alerting systems

Nice to have

  • Experience with GPU infrastructure or AI/ML platforms
  • Experience improving reliability in high-growth or large scale environments
  • Familiarity with GPU observability tooling
  • Experience with Infrastructure as Code
  • Experience working in startup environments

Skills

  • Our platform enables teams to move from experimentation to deployment with flexibility across cloud, on-prem, and hybrid environments.
  • Defining and enforcing reliability standards across engineering
  • Designing incident response processes and improving recovery times
  • Building observability systems and reliability tooling
  • Driving SLO adoption and production readiness reviews
  • Increase platform uptime and reduce incident frequency and duration
  • Establish and operationalize SLIs/SLOs across services
  • Improve MTTR through better tooling, automation, and runbooks
  • Strengthen production readiness standards
  • Drive long-term systemic reliability improvements

Compensation

  • The competitive base pay for this position ranges from $150,000- $200,000 usd.

Benefits

  • Meaningful equity in a fast-growing company- everyone on the team receives stock options - your impact drives our growth, and you share in the upside.
  • Generous medical, dental & vision plans
  • Flexible PTO- take the time you need to recharge
  • Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.
  • Improve visibility into GPU performance and distributed systems health

Equal opportunity

  • As an equal opportunity employer, RunPod is committed to creating an inclusive workforce at every level.

This listing is sourced directly from Runpod's careers page and normalized into a canonical job model.