Qumulo
Site Reliability Engineer (SRE) - Hybrid Cloud Storage
Seattle (hybrid)
Sponsorship not specified$140k-$210kDetected 21 days ago
PythonAWSGCPAzureKubernetesTerraformAnsibleJenkinsLinuxPrometheusGrafanaSite Reliability EngineeringCustomer SupportTest AutomationFirewallBioinformaticsRadiology
About the role
- You'll be one of the first hires on a team with a single mandate: find out how Qumulo breaks before our customers do.
- This is a test-centric SRE role for an engineer who thinks like a breaker.
- You'll automate the testing our principal engineers run by hand today, and decide what gets tested, how often, and why.
Responsibilities
- You'll help build this function from the ground up, including where we set the quality bar and which builds are good enough to ship.
- Troubleshoot build and test failures across VM instances and Qumulo-qualified hardware, from compile-time errors to integration failures
- At Qumulo, we are building an open and collaborative culture where people can do their best work with customers as our magnetic field.
- As part of our culture we believe diversity drives innovation.
Requirements
- Experience with orchestration tools (Ansible, Terraform), containers, and Kubernetes
Nice to have
- Automate the manual, repetitive testing our principal engineers run by hand today, using Python and our in-house frameworks on Jenkins and Argo
- Read cluster output and C error logs to tell a test problem from an infrastructure problem from a real bug
- Set up monitoring and alerting so problems surface early (we use OpenMetrics, Grafana, InfluxDB, and Prometheus alongside home-grown tooling)
- Help set the quality bar for releases, including a real say in what ships
- Strong programming ability in C.
- Experience with Qumulo's distributed file system, or parallel filesystems, would be a major plus
- Solid understanding of networks (routing, firewalls, security inspection devices, switch configuration) a plus
- Storage (IOPS, Latency, read/write patterns) or protocol experience (NFS, SMB, S3, ) a strong plus
Skills
- Hands-on experience across both on-premises infrastructure and cloud (AWS, GCP, or Azure), with a real grasp of where each one's limits are
Compensation
- 3+ years building and operating automated testing, validation, and/or certification for complex software systems
- Strong programming ability in C. Experience with Qumulo's distributed file system, or parallel filesystems, would be a major plus
- A real breaker's instinct. You go looking for edge cases and ask "what happens if I do this?" before anyone asks you to
- A track record of building tests yourself, not just running test plans handed to you
- Hands-on experience across both on-premises infrastructure and cloud (AWS, GCP, or Azure), with a real grasp of where each one's limits are
- Working fluency in Linux (we run Ubuntu) and Python
Benefits
- Pre-IPO stock options
- Flexible time-off policy
- HSA and PPO health insurance options
- Dental and Vision insurance
Company info
- We act as owners, we share by default, we are data driven and experimental and as an inclusive workplace, we encourage and celebrate multiple points of view.
- Design and operationalize testing for new features: work out how customers will actually use them, how to scale-test them, and how to break them
- With more than 1,100 customers and exabytes of data under management, Qumulo powers mission-critical workloads anywhere real-time access to massive file datasets is non-negotiable.
Equal opportunity
- Employer:
- Qumulo is an Equal Opportunity Employer.
- Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, disability, military status, national origin, or any other characteristic protected under federal, state, or applicable local law.
- Equal Opportunity Employer:
This listing is sourced directly from Qumulo's careers page and normalized into a canonical job model.