Backblaze External Website

Backblaze External Website

Sr. Site Reliability Engineer

Remote - US · Senior · Full-time

Sponsorship not specified$150k-$200kDetected 112 days ago
PythonGoBashDistributed SystemsAWSGCPAzureCloud PlatformsDockerKubernetesTerraformAnsibleCI/CDJenkinsLinuxPrometheusGrafanaSite Reliability EngineeringIncident ResponseCustomer SuccessSystems EngineeringLeadershipCollaborationProblem Solving

About the role

  • We are seeking a Senior Site Reliability Engineer (SRE) to help ensure the stability, scalability, and reliability of our services and infrastructure.
  • Mentor others and act as a subject matter expert in following and evolving established ITIL/OSS processes (incident, change, problem, and capacity management).
  • Be a leading voice in promoting and embedding reliability-focused practices within development and operations teams.

Responsibilities

  • Own and drive the availability, durability, and performance of critical services across all production environments.
  • Lead and champion complex projects from problem discovery through complete, cross-functional resolution, demonstrating high-level technical ownership.
  • Lead critical incident response and post-incident reviews, translating findings into strategic, long-term service improvements and architectural changes.
  • Design and architect scalable automation solutions to eliminate toil and improve the efficiency of operational tasks across the entire platform.
  • Drive the strategic direction of monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint, ELK), and integrate them for comprehensive observability.
  • Build, maintain, and secure advanced CI/CD pipelines, configuration management, and complex infrastructure as code solutions (Terraform, Ansible, Jenkins).
  • Write production-grade code (Bash, Python, Go, etc.) to develop new reliability tools and enhance existing systems.
  • Act as a principal partner to engineering, product, and operations teams, consulting on resilient system design, architecture, and operation.
  • Lead and formalize the Production Readiness Review (PRR) process, ensuring robust operational handoff for all new services and features.
  • Fertility treatment and support

Requirements

  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience).
  • 8+ years of progressive experience in site reliability, systems engineering, or operations.
  • Extensive experience designing, scaling, and operating large-scale, production-grade distributed systems.
  • Expert knowledge of incident response methodologies and operational best practices.
  • Proven experience designing and operating container orchestration (Kubernetes, Docker) and microservices concepts required.
  • Expert experience with Hashicorp products (Nomad, Vault, Terraform) in a production environment.
  • Significant experience in a SaaS, service provider, or hyper-scale distributed systems environment.
  • Advanced experience with cloud platforms (AWS, GCP, or Azure) in a production setting.
  • Education & Experience
  • Technical Skills
  • Expert-level Linux systems administration and advanced troubleshooting skills.
  • Lead security-minded operations, focusing on system-wide patching, hardening, and proactive vulnerability identification.
  • Deep mastery of service reliability concepts, including advanced monitoring, complex alerting strategy, leading incident response, and in-depth root cause analysis.

Nice to have

  • Advanced proficiency in at least one modern scripting/programming language (Python or Go strongly preferred).
  • Preferred Attributes

Skills

  • About Backblaze
  • But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us.

Compensation

  • Competitive compensation and 401K

Benefits

  • Healthcare for family, including dental and vision
  • RSU grants for full-time employees
  • Flexible vacation policy
  • Maternity & paternity leave
  • MacBook Pro to use for work, plus a generous stipend to personalize your workstation
  • Childcare bonus (human children only)
  • Learning & development program
  • Commuter benefits
  • Define, establish, and enforce service health standards, including working with engineering leadership to implement SLIs, SLOs, and error budget policies for multiple services.

This listing is sourced directly from Backblaze External Website's careers page and normalized into a canonical job model.