Runpod
Site Reliability Engineer
Remote, USA
Sponsorship not specified$150k-$200kDetected 60 days ago
PythonGoBashDistributed SystemsCI/CDLinuxPrometheusGrafanaSite Reliability EngineeringMachine LearningA/B TestingIncident ResponseLeadershipCommunicationProblem Solving
About the role
- With over 500,000 developers worldwide and an annual recurring revenue run rate exceeding $120M, Runpod operates at the intersection of developer velocity and production-scale AI.
- We value proactive problem solving, automation-first thinking, and strong ownership of production systems.
- This role blends software engineering with production operations.
Responsibilities
- Define and implement SLIs/SLOs for critical services
- Lead incident response and coordinate cross-team mitigation efforts
- Perform production readiness reviews for new services and features
- Identify systemic risks and drive preventative improvements
- Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
- Build internal tooling for reliability tracking and reporting
- Build tools and scripts (Python, Go, Bash) to eliminate manual processes
- Partner with engineering teams to improve system resilience
- You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company.
Requirements
- 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
- Strong Linux systems and Networking expertise
- Experience managing containerized production systems
- Strong understanding of distributed systems and failure modes
- Experience defining and managing SLIs/SLOs
- Proven incident response and postmortem leadership experience
- Experience with monitoring and alerting systems
Nice to have
- Experience with GPU infrastructure or AI/ML platforms
- Experience improving reliability in high-growth or large scale environments
- Familiarity with GPU observability tooling
- Experience with Infrastructure as Code
- Experience working in startup environments
Skills
- Our platform enables teams to move from experimentation to deployment with flexibility across cloud, on-prem, and hybrid environments.
- Defining and enforcing reliability standards across engineering
- Designing incident response processes and improving recovery times
- Building observability systems and reliability tooling
- Driving SLO adoption and production readiness reviews
- Increase platform uptime and reduce incident frequency and duration
- Establish and operationalize SLIs/SLOs across services
- Improve MTTR through better tooling, automation, and runbooks
- Strengthen production readiness standards
- Drive long-term systemic reliability improvements
Compensation
- The competitive base pay for this position ranges from $150,000- $200,000 usd.
Benefits
- Meaningful equity in a fast-growing company- everyone on the team receives stock options - your impact drives our growth, and you share in the upside.
- Generous medical, dental & vision plans
- Flexible PTO- take the time you need to recharge
- Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.
- Improve visibility into GPU performance and distributed systems health
Equal opportunity
- As an equal opportunity employer, RunPod is committed to creating an inclusive workforce at every level.
This listing is sourced directly from Runpod's careers page and normalized into a canonical job model.