Stuut

Stuut

Lead Site Reliability Engineer

San Francisco · Full-time

Sponsorship not specifiedDetected 3 days ago
TypeScriptPythonVue.jsFastAPIDistributed SystemsPostgreSQLAWSCloud PlatformsDockerKubernetesCI/CDSite Reliability EngineeringIncident ResponseLeadershipAccounts Payable

About the role

  • You'll turn strong reliability engineering into real customer trust - creating the guardrails that let product and engineering move fast with confidence.
  • This is a hands-on technical leadership role for an engineer who excels at designing reliable distributed systems, influencing engineering practices, and leading high-impact reliability initiatives across teams.

Responsibilities

  • Build & Scale Reliable Infrastructure: architect and maintain resilient, scalable cloud infrastructure across AWS and Kubernetes, ensuring systems are secure, fault-tolerant, and cost-effective.
  • Own Observability & Monitoring: design and evolve monitoring, alerting, and logging systems that provide clear, actionable signals across services and environments.
  • Lead Incident Response & Postmortems: own incident management practices, lead major incident response, and drive blameless postmortems that result in meaningful system improvements.
  • Improve System Resilience: identify reliability risks and lead efforts around redundancy, failover, capacity planning, and graceful degradation.
  • Optimize CI/CD & Deployment Reliability: partner with engineering teams to ensure deployments are safe, observable, and reversible
  • Partner with Product & Engineering Teams: collaborate early in the development lifecycle to influence system design, scalability, and reliability tradeoffs.
  • Reduce Toil & Improve Developer Experience: automate operational tasks, improve runbooks, and build tooling that reduces manual work and accelerates safe execution.
  • Drive Root Cause Resolution: guide teams through deep debugging of reliability issues, ensuring fixes address underlying causes rather than symptoms.
  • Mentor & Level Up the Team: coach engineers on reliability principles, incident handling, infrastructure design, and operational best practices.

Requirements

  • Have 7+ years of experience in site reliability engineering, infrastructure engineering, or backend software engineering.
  • Have a deep experience with AWS, Kubernetes (EKS), Docker, and cloud-native architectures.

Compensation

  • Top-of-market salary and equity package

Benefits

  • define the long-term vision for site reliability, including SLOs/SLIs, error budgets, availability targets, and operational standards.
  • Benefits (for U.S.-based full-time employees)
  • Medical, dental & vision insurance coverage for you
  • Parental Leave

Company info

  • promote reliability-first thinking, strong operational hygiene, and shared ownership of production systems across engineering.
  • Have designed and operated highly available, production-grade systems supporting rapid product iteration.
  • Are fluent in Python and/or TypeScript, and comfortable building automation and tooling to support reliability goals.
  • Have implemented and evolved observability stacks (metrics, logs, traces) and know how to create high-signal alerting.
  • Understand how to design, measure, and enforce SLOs, SLIs, and error budgets.
  • Have supported systems built with modern stacks such as FastAPI, Vue.js, PostgreSQL (RDS), and event-driven architectures.
  • Have improved reliability and operational maturity in environments using CI/CD pipelines, infrastructure as code, and modern deployment workflows.
  • Can balance reliability, velocity, and cost - making pragmatic tradeoffs that serve customers and the business.
  • Enjoy collaborating across Product, Backend, Frontend, and Infrastructure teams to improve system health.
  • Thrive in a role that blends deep technical execution, system design, and leadership influence in a fast-moving environment.

This listing is sourced directly from Stuut's careers page and normalized into a canonical job model.