Backblaze External Website
Sr. Site Reliability Engineer
Remote - US · Senior · Full-time
Sponsorship not specified$150k-$200kDetected 112 days ago
PythonGoBashDistributed SystemsAWSGCPAzureCloud PlatformsDockerKubernetesTerraformAnsibleCI/CDJenkinsLinuxPrometheusGrafanaSite Reliability EngineeringIncident ResponseCustomer SuccessSystems EngineeringLeadershipCollaborationProblem Solving
About the role
- We are seeking a Senior Site Reliability Engineer (SRE) to help ensure the stability, scalability, and reliability of our services and infrastructure.
- Mentor others and act as a subject matter expert in following and evolving established ITIL/OSS processes (incident, change, problem, and capacity management).
- Be a leading voice in promoting and embedding reliability-focused practices within development and operations teams.
Responsibilities
- Own and drive the availability, durability, and performance of critical services across all production environments.
- Lead and champion complex projects from problem discovery through complete, cross-functional resolution, demonstrating high-level technical ownership.
- Lead critical incident response and post-incident reviews, translating findings into strategic, long-term service improvements and architectural changes.
- Design and architect scalable automation solutions to eliminate toil and improve the efficiency of operational tasks across the entire platform.
- Drive the strategic direction of monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint, ELK), and integrate them for comprehensive observability.
- Build, maintain, and secure advanced CI/CD pipelines, configuration management, and complex infrastructure as code solutions (Terraform, Ansible, Jenkins).
- Write production-grade code (Bash, Python, Go, etc.) to develop new reliability tools and enhance existing systems.
- Act as a principal partner to engineering, product, and operations teams, consulting on resilient system design, architecture, and operation.
- Lead and formalize the Production Readiness Review (PRR) process, ensuring robust operational handoff for all new services and features.
- Fertility treatment and support
Requirements
- Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience).
- 8+ years of progressive experience in site reliability, systems engineering, or operations.
- Extensive experience designing, scaling, and operating large-scale, production-grade distributed systems.
- Expert knowledge of incident response methodologies and operational best practices.
- Proven experience designing and operating container orchestration (Kubernetes, Docker) and microservices concepts required.
- Expert experience with Hashicorp products (Nomad, Vault, Terraform) in a production environment.
- Significant experience in a SaaS, service provider, or hyper-scale distributed systems environment.
- Advanced experience with cloud platforms (AWS, GCP, or Azure) in a production setting.
- Education & Experience
- Technical Skills
- Expert-level Linux systems administration and advanced troubleshooting skills.
- Lead security-minded operations, focusing on system-wide patching, hardening, and proactive vulnerability identification.
- Deep mastery of service reliability concepts, including advanced monitoring, complex alerting strategy, leading incident response, and in-depth root cause analysis.
Nice to have
- Advanced proficiency in at least one modern scripting/programming language (Python or Go strongly preferred).
- Preferred Attributes
Skills
- About Backblaze
- But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us.
Compensation
- Competitive compensation and 401K
Benefits
- Healthcare for family, including dental and vision
- RSU grants for full-time employees
- Flexible vacation policy
- Maternity & paternity leave
- MacBook Pro to use for work, plus a generous stipend to personalize your workstation
- Childcare bonus (human children only)
- Learning & development program
- Commuter benefits
- Define, establish, and enforce service health standards, including working with engineering leadership to implement SLIs, SLOs, and error budget policies for multiple services.
Apply directly at Backblaze External Website →Create a free account for alerts like thisView Backblaze External Website immigration profile
This listing is sourced directly from Backblaze External Website's careers page and normalized into a canonical job model.