GRAIL
Staff Site Reliability Engineer (SRE)
Menlo Park, CA · Staff+ · Full-time
Sponsorship not specified$169k-$224kDetected 91 days ago
PythonGoBashPowerShellDistributed SystemsGitBigQuerySnowflakeAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleCI/CDGitHub ActionsJenkinsPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringPlatform Engineering
About the role
- GRAIL is headquartered in the bay area of California, with locations in Washington, D.C., North Carolina, and the United Kingdom.
- It is supported by leading global investors and pharmaceutical, technology, and healthcare companies.
- This is a highly technical, high-impact role combining hands-on engineering with cross-functional influence and mentorship.
Responsibilities
- Design, build, and operate highly available, fault-tolerant cloud infrastructure across AWS, GCP, and/or Azure
- Architect and maintain scalable CI/CD pipelines and deployment frameworks for enterprise-grade software delivery
- Lead infrastructure-as-code adoption and maturity using tools such as Terraform, CloudFormation, and Ansible
- Own Kubernetes reliability across multi-cluster environments, including upgrades, scaling, and workload lifecycle management
- Lead incident response for critical outages, drive root cause analysis, and implement preventative improvements
- Optimize infrastructure for cost, performance, and scalability, partnering closely with engineering and finance stakeholders
- Partner cross-functionally with engineering, data, QA, security, and IT teams to design resilient systems
- Mentor engineers and contribute to technical leadership through design reviews, standards, and knowledge sharing
- Standardize observability using modern tooling and implement an SLO/SLI framework adopted across multiple product teams, including defined SLAs for critical data systems
- Define, document, and drive adoption of engineering standards, best practices, and operational guidelines across platform and product teams
Requirements
- BS in Computer Science, Engineering, or related field, or equivalent experience
- 8+ years of experience in Site Reliability Engineering, DevOps, or platform engineering
- Strong hands-on experience with at least one major cloud platform (AWS, GCP, or Azure)
- Experience implementing infrastructure-as-code solutions (Terraform, CloudFormation, or similar)
- Experience designing and operating CI/CD pipelines (e.g., GitLab CI, GitHub Actions, Jenkins)
- Hands-on experience with Kubernetes and containerized systems in production environments
- Proficiency in scripting or programming for automation (e.g., Python, Go, Bash, or PowerShell)
- Experience with observability and monitoring tools (e.g., Prometheus, Grafana, OpenTelemetry, Datadog)
- Strong understanding of networking, security, and distributed systems fundamentals
- Experience working in regulated environments and familiarity with frameworks such as ISO 27001, NIST, SOC 2, or HIPAA
Nice to have
- 10+ years of experience in SRE, DevOps, or infrastructure engineering
- Experience operating multi-cluster Kubernetes environments (e.g., EKS, GKE) at scale
- Familiarity with GitOps practices (e.g., ArgoCD, Flux)
- Experience with data platforms and pipelines (e.g., Kafka, Airflow, Spark, Snowflake, BigQuery)
- Experience implementing SLO/SLI frameworks and reliability practices across multiple teams
- Strong background in cloud security, including IAM, zero-trust architecture, and secrets management
- Experience with compliance-as-code and security tooling (e.g., OPA, Snyk, Checkov)
- Exposure to AI/ML or large-scale data infrastructure workloads
Compensation
- The expected, full-time, annual base pay scale for this position is $169K - $224K.
Benefits
- Conduct a comprehensive assessment of the current infrastructure, drive infrastructure-as-code adoption to 95%+ across critical systems, and establish clear health and reliability baselines for the Kubernetes platform
This listing is sourced directly from GRAIL's careers page and normalized into a canonical job model.