Backblaze External Website
Cluster & Systems Capacity Engineer
Remote - US · Full-time
Sponsorship not specified$123k-$175kDetected 22 days ago
PythonDistributed SystemsSQLSnowflakePrometheusGrafanaSite Reliability EngineeringData AnalysisData ScienceStatisticsForecastingExcelOnboardingCustomer SuccessNetwork EngineeringCommunication
About the role
- This role ensures that Backblaze's storage clusters, compute systems, and network infrastructure scale reliably, cost-efficiently, and ahead of demand.
- This is a high-impact role within Cloud Operations, directly contributing to service availability, durability, performance, margin optimization, and long-term platform scalability.
Responsibilities
- Develop and maintain short, medium, and long-term capacity demand and hardware deployment forecasts across storage, compute, and network domains within the platform
- Build predictive models that translate business demand signals into infrastructure requirements using historical utilization, growth trends, product sales plans, hardware lifecycle roadmaps, and other key business inputs
- Partner with Infrastructure, Production, and Network Engineering teams to align capacity plans with system design and scaling initiatives
- Adjust deployment plans and recommended configurations in real-time to maintain adequate headroom and system stability in support of delivering a world-class customer experience
- Partner with service and platform owners to develop headroom and live buffer policies, optimize hardware BoMs, leverage virtualized orchestration, and reduce product cost
- Support strategic optimization initiatives across infrastructure investments, engineering development, and operations processes, contributing to long-term infrastructure strategy and capital planning
- Lead efforts to evaluate, procure, and provision requests for new or additional hardware, working with Systems and Network Engineering, SRE, NOC, and Data Center Operations teams to identify and deliver optimal solutions
- Maintain alignment with Product and Sales to support customer onboarding, growth, and demand variability
- Fertility treatment and support
- Culture that supports a healthy work-life balance
Requirements
- Bachelor's degree in Computer Science, Engineering, Mathematics, Data Science, Information Systems, Statistics or a related, technical field (or equivalent experience).
- 3-6+ years of experience in Site Reliability Engineering, Infrastructure Capacity Planning, Systems/Infrastructure Engineering, Production Engineering, Data Center Operations or similar Cloud Operations role
- Background in capacity modeling, performance analysis, scenario modeling, and/or infrastructure cost optimization, with an ability to quantify ideas within financial frameworks and forecasts.
- Proficiency in database and data analysis tools (preferably Snowflake, Metabase, Grafana, Python, SQL, Prometheus, Victoria Metrics, and Excel/Google Sheets)
- Excellent communication and documentation skills, with the ability to share knowledge and explain concepts accurately and concisely
Skills
- About Backblaze
- But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us.
Compensation
- Competitive compensation and 401K
Benefits
- Healthcare for family, including dental and vision
- RSU grants for full-time employees
- Flexible vacation policy
- Maternity & paternity leave
- MacBook Pro to use for work, plus a generous stipend to personalize your workstation
- Childcare bonus (human children only)
- Learning & development program
- Commuter benefits
- Learning, developing, and growing are key parts of our culture.
Company info
- At this point, we hope you're feeling excited about the job description you're reading.
- We're eager to meet people who believe in our mission and can contribute to our team in various ways.
Apply directly at Backblaze External Website →Create a free account for alerts like thisView Backblaze External Website immigration profile
This listing is sourced directly from Backblaze External Website's careers page and normalized into a canonical job model.