Drata

Drata

Senior Site Reliability Engineer

Hybrid - San Francisco · Senior

Sponsorship not specified$167k-$226kDetected 85 days ago
JavaScriptPythonBashNode.jsGitMySQLAWSDockerKubernetesTerraformCI/CDGitHub ActionsLinuxDatadogSite Reliability EngineeringRESTMachine LearningLLMsIncident ResponseComplianceCollaboration

About the role

  • Drata's SRE team operates as both a central engineering function and an embedded reliability practice.
  • This is a highly technical role at the intersection of software engineering and systems engineering.
  • Automation is a core value, and nowhere is that more visible than in how we approach reliability.

Responsibilities

  • Lead Production Readiness Reviews (PRRs) before new services launch, with the authority to flag gaps and gate launches when critical reliability standards aren't met
  • Partner with product engineering leads and staff engineers to define SLOs and SLIs for critical services, turning reliability from a vague goal into a measurable commitment
  • Build reusable artifacts - SLO templates, observability checklists, alerting standards, reference dashboards - that raise the reliability floor across the team, not just the services you touch directly
  • When an engineer needs something, your priority is: automate it so anyone can do it → document it so the team can self-serve → execute it manually only as a last resort.
  • Build and maintain Datadog monitors, dashboards, and alert routing - enforcing infrastructure-as-code standards via Terraform so those resources are owned, versioned, and auditable

Requirements

  • 6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building and maintaining scalable, resilient services
  • You are the reliability expert for your aligned product team.

Nice to have

  • Experience with AIOps - using AI/ML-based tooling for anomaly detection, predictive alerting, or automated incident triage
  • Familiarity with the reliability characteristics of AI/ML-backed services (e.g., LLM inference latency, non-determinism, prompt pipeline observability)
  • Experience with the JavaScript/Node.js ecosystem
  • Certified Kubernetes Administrator (CKA) certification
  • Familiarity with compliance frameworks like SOC 2, ISO 27001, or NIST
  • AI Experience (required - at least one of the following):
  • Hands-on experience using AI-assisted development tools (e.g., GitHub Copilot, Cursor, or similar) to accelerate automation, scripting, or infrastructure work
  • Demonstrated use of AI/AIOps capabilities for reliability tasks - anomaly detection, incident triage, runbook generation, or alert noise reduction

Skills

  • Hands-on experience with Datadog for monitoring, alerting, dashboards, SLO tracking, and distributed tracing
  • Experience building software systems as a software engineer
  • Experience with CI/CD pipeline automation, specifically GitHub Actions
  • Experience with disaster recovery practices and incident management
  • Experience with container orchestration and deployment technologies including AWS ECS Fargate and/or Kubernetes
  • Experience working with relational databases (MySQL proficiency is a plus)
  • Ability to take ownership of problems and act on them independently in a constantly evolving environment
  • Terraform, Docker, Git, and Linux
  • Our infrastructure runs on AWS across multiple accounts, defined entirely in Terraform.

Compensation

  • A variety of factors are considered when determining someone's leveling and compensation-including a candidate's professional background and experience.
  • This role will receive a competitive base salary, benefits, and stock, typically in the form of Restricted Stock Units (RSUs).
  • $166,900 - $225,900.

Benefits

  • We provide stock equity to ensure that as the company grows, you share directly in that success.
  • Equity gives every employee a sense of ownership and the opportunity to celebrate our wins together-because your contributions don't just support our progress; they help drive our collective success.
  • We want to support you in life's most important moments, so we offer a paid Parental Leave policy, after six months of employment.
  • Employees also receive access to Kindbody fertility and family-building benefits and dedicated leave specialists who help guide you through the entire process.
  • Beyond your product team, you contribute to cross-cutting infrastructure, tooling, and standards that benefit every team at Drata.
  • Drata offers a flexible vacation policy, paid holidays, and other perks to recharge.

Company info

  • Hear the Voice of the Team https://drata.com/about/life-at-drata: Explore our "Life at Drata" page for employee testimonials on our collaborative and the growth opportunities available.
  • Experience the Impact https://www.greatplacetowork.com/certified-company/7044563: See why we are consistently recognized on Fortune's Best Workplaces lists.
  • LinkedIn https://www.linkedin.com/company/drata/posts/?feedView=all - follow us for company updates, employee stories, and career news.
  • A comprehensive suite of financial benefits, including a 401(k) plan, company-paid life and disability insurance, tax-advantaged spending accounts, and a range of discounted voluntary offerings to help you customize and strengthen your overall financial position.

This listing is sourced directly from Drata's careers page and normalized into a canonical job model.