Drata
Senior Site Reliability Engineer
Hybrid - San Francisco · Senior
Sponsorship not specified$167k-$226kDetected 85 days ago
JavaScriptPythonBashNode.jsGitMySQLAWSDockerKubernetesTerraformCI/CDGitHub ActionsLinuxDatadogSite Reliability EngineeringRESTMachine LearningLLMsIncident ResponseComplianceCollaboration
About the role
- Drata's SRE team operates as both a central engineering function and an embedded reliability practice.
- This is a highly technical role at the intersection of software engineering and systems engineering.
- Automation is a core value, and nowhere is that more visible than in how we approach reliability.
Responsibilities
- Lead Production Readiness Reviews (PRRs) before new services launch, with the authority to flag gaps and gate launches when critical reliability standards aren't met
- Partner with product engineering leads and staff engineers to define SLOs and SLIs for critical services, turning reliability from a vague goal into a measurable commitment
- Build reusable artifacts - SLO templates, observability checklists, alerting standards, reference dashboards - that raise the reliability floor across the team, not just the services you touch directly
- When an engineer needs something, your priority is: automate it so anyone can do it → document it so the team can self-serve → execute it manually only as a last resort.
- Build and maintain Datadog monitors, dashboards, and alert routing - enforcing infrastructure-as-code standards via Terraform so those resources are owned, versioned, and auditable
Requirements
- 6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building and maintaining scalable, resilient services
- You are the reliability expert for your aligned product team.
Nice to have
- Experience with AIOps - using AI/ML-based tooling for anomaly detection, predictive alerting, or automated incident triage
- Familiarity with the reliability characteristics of AI/ML-backed services (e.g., LLM inference latency, non-determinism, prompt pipeline observability)
- Experience with the JavaScript/Node.js ecosystem
- Certified Kubernetes Administrator (CKA) certification
- Familiarity with compliance frameworks like SOC 2, ISO 27001, or NIST
- AI Experience (required - at least one of the following):
- Hands-on experience using AI-assisted development tools (e.g., GitHub Copilot, Cursor, or similar) to accelerate automation, scripting, or infrastructure work
- Demonstrated use of AI/AIOps capabilities for reliability tasks - anomaly detection, incident triage, runbook generation, or alert noise reduction
Skills
- Hands-on experience with Datadog for monitoring, alerting, dashboards, SLO tracking, and distributed tracing
- Experience building software systems as a software engineer
- Experience with CI/CD pipeline automation, specifically GitHub Actions
- Experience with disaster recovery practices and incident management
- Experience with container orchestration and deployment technologies including AWS ECS Fargate and/or Kubernetes
- Experience working with relational databases (MySQL proficiency is a plus)
- Ability to take ownership of problems and act on them independently in a constantly evolving environment
- Terraform, Docker, Git, and Linux
- Our infrastructure runs on AWS across multiple accounts, defined entirely in Terraform.
Compensation
- A variety of factors are considered when determining someone's leveling and compensation-including a candidate's professional background and experience.
- This role will receive a competitive base salary, benefits, and stock, typically in the form of Restricted Stock Units (RSUs).
- $166,900 - $225,900.
Benefits
- We provide stock equity to ensure that as the company grows, you share directly in that success.
- Equity gives every employee a sense of ownership and the opportunity to celebrate our wins together-because your contributions don't just support our progress; they help drive our collective success.
- We want to support you in life's most important moments, so we offer a paid Parental Leave policy, after six months of employment.
- Employees also receive access to Kindbody fertility and family-building benefits and dedicated leave specialists who help guide you through the entire process.
- Beyond your product team, you contribute to cross-cutting infrastructure, tooling, and standards that benefit every team at Drata.
- Drata offers a flexible vacation policy, paid holidays, and other perks to recharge.
Company info
- Hear the Voice of the Team https://drata.com/about/life-at-drata: Explore our "Life at Drata" page for employee testimonials on our collaborative and the growth opportunities available.
- Experience the Impact https://www.greatplacetowork.com/certified-company/7044563: See why we are consistently recognized on Fortune's Best Workplaces lists.
- LinkedIn https://www.linkedin.com/company/drata/posts/?feedView=all - follow us for company updates, employee stories, and career news.
- A comprehensive suite of financial benefits, including a 401(k) plan, company-paid life and disability insurance, tax-advantaged spending accounts, and a range of discounted voluntary offerings to help you customize and strengthen your overall financial position.
This listing is sourced directly from Drata's careers page and normalized into a canonical job model.