GRAIL

GRAIL

Staff Site Reliability Engineer (SRE)

Menlo Park, CA · Staff+ · Full-time

Sponsorship not specified$169k-$224kDetected 91 days ago
PythonGoBashPowerShellDistributed SystemsGitBigQuerySnowflakeAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleCI/CDGitHub ActionsJenkinsPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringPlatform Engineering

About the role

  • GRAIL is headquartered in the bay area of California, with locations in Washington, D.C., North Carolina, and the United Kingdom.
  • It is supported by leading global investors and pharmaceutical, technology, and healthcare companies.
  • This is a highly technical, high-impact role combining hands-on engineering with cross-functional influence and mentorship.

Responsibilities

  • Design, build, and operate highly available, fault-tolerant cloud infrastructure across AWS, GCP, and/or Azure
  • Architect and maintain scalable CI/CD pipelines and deployment frameworks for enterprise-grade software delivery
  • Lead infrastructure-as-code adoption and maturity using tools such as Terraform, CloudFormation, and Ansible
  • Own Kubernetes reliability across multi-cluster environments, including upgrades, scaling, and workload lifecycle management
  • Lead incident response for critical outages, drive root cause analysis, and implement preventative improvements
  • Optimize infrastructure for cost, performance, and scalability, partnering closely with engineering and finance stakeholders
  • Partner cross-functionally with engineering, data, QA, security, and IT teams to design resilient systems
  • Mentor engineers and contribute to technical leadership through design reviews, standards, and knowledge sharing
  • Standardize observability using modern tooling and implement an SLO/SLI framework adopted across multiple product teams, including defined SLAs for critical data systems
  • Define, document, and drive adoption of engineering standards, best practices, and operational guidelines across platform and product teams

Requirements

  • BS in Computer Science, Engineering, or related field, or equivalent experience
  • 8+ years of experience in Site Reliability Engineering, DevOps, or platform engineering
  • Strong hands-on experience with at least one major cloud platform (AWS, GCP, or Azure)
  • Experience implementing infrastructure-as-code solutions (Terraform, CloudFormation, or similar)
  • Experience designing and operating CI/CD pipelines (e.g., GitLab CI, GitHub Actions, Jenkins)
  • Hands-on experience with Kubernetes and containerized systems in production environments
  • Proficiency in scripting or programming for automation (e.g., Python, Go, Bash, or PowerShell)
  • Experience with observability and monitoring tools (e.g., Prometheus, Grafana, OpenTelemetry, Datadog)
  • Strong understanding of networking, security, and distributed systems fundamentals
  • Experience working in regulated environments and familiarity with frameworks such as ISO 27001, NIST, SOC 2, or HIPAA

Nice to have

  • 10+ years of experience in SRE, DevOps, or infrastructure engineering
  • Experience operating multi-cluster Kubernetes environments (e.g., EKS, GKE) at scale
  • Familiarity with GitOps practices (e.g., ArgoCD, Flux)
  • Experience with data platforms and pipelines (e.g., Kafka, Airflow, Spark, Snowflake, BigQuery)
  • Experience implementing SLO/SLI frameworks and reliability practices across multiple teams
  • Strong background in cloud security, including IAM, zero-trust architecture, and secrets management
  • Experience with compliance-as-code and security tooling (e.g., OPA, Snyk, Checkov)
  • Exposure to AI/ML or large-scale data infrastructure workloads

Compensation

  • The expected, full-time, annual base pay scale for this position is $169K - $224K.

Benefits

  • Conduct a comprehensive assessment of the current infrastructure, drive infrastructure-as-code adoption to 95%+ across critical systems, and establish clear health and reliability baselines for the Kubernetes platform

This listing is sourced directly from GRAIL's careers page and normalized into a canonical job model.