LeanData
Senior Site Reliability Engineer
Santa Clara · Senior · Full-time
Sponsorship not specifiedDetected 165 days ago
PythonC#BashAngularNode.jsRedisAWSCloud PlatformsTerraformCI/CDDevOpsSite Reliability EngineeringCybersecurityIncident ResponseSalesforceSystems EngineeringFirewallLeadership
About the role
- You will be the primary authority on reliability, performance, and infrastructure security.
- Please note: This is a hybrid role based in our Santa Clara, CA office, with an in-office schedule of two days per week - Monday and Wednesday.
Responsibilities
- Architectural Modernization: Lead the design and implementation of a scalable, "Cloud-First" AWS architecture. You will drive the transition toward fully automated, state-of-the-art Infrastructure as Code (Terraform).
- High Availability & Resilience: Design and implement robust Disaster Recovery (DR) and Business Continuity plans, moving our services toward a zero-downtime deployment model.
- Performance & Capacity Engineering: Own the strategy for capacity planning and autoscaling. You will optimize our compute resources (EC2, Lambda) to handle bursty traffic patterns with precision and cost-efficiency.
- Streamlined CI/CD: Partner with feature teams to refine Change Management and CI/CD pipelines, ensuring code moves from "commit" to "production" safely and predictably.
- Architectural Modernization: Lead the design and implementation of a scalable, "Cloud-First" AWS architecture.
- You will drive the transition toward fully automated, state-of-the-art Infrastructure as Code (Terraform).
- Design and implement robust Disaster Recovery (DR) and Business Continuity plans, moving our services toward a zero-downtime deployment model.
- Own the strategy for capacity planning and autoscaling.
- You will optimize our compute resources (EC2, Lambda) to handle bursty traffic patterns with precision and cost-efficiency.
- Partner with feature teams to refine Change Management and CI/CD pipelines, ensuring code moves from "commit" to "production" safely and predictably.
Requirements
- Experienced Architect: 5+ years of experience in SRE, DevOps, or Systems Engineering, with a proven track record of managing complex AWS environments.
- Proven Incident Commander: You demonstrate calm, decisive leadership during high-pressure outages.
- You have extensive experience running blameless postmortems and, crucially, driving the remediation work needed to prevent recurrence.
- You have deep experience with Terraform and a "Code-First" approach to infrastructure.
- Strategic Problem Solver: You can look at a complex, "needs-based" architecture and formulate a clear, prioritized roadmap to move it toward industry best practices.
- 5+ years of experience in SRE, DevOps, or Systems Engineering, with a proven track record of managing complex AWS environments.
- A Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent professional experience).
Skills
- Terraform, Redis/Elasticache, Shell Scripting, NPM/PM2.
- LeanData helps the world's fastest-growing companies automate, simplify, and accelerate revenue.
- Harden our network architecture and application security posture, including WAF management and secure service-to-service communication.
- AWS (EC2, Lambda, SQS, SNS, ALB, API Gateway, S3, WAF).
Benefits
- Stock options in LeanData for all full-time employees
Company info
- Define our monitoring and alerting philosophy using New Relic for deep APM and system insights.
- Partner this with IncidentIO to ensure we catch and resolve issues before they impact customers.
- Advanced Observability: Define our monitoring and alerting philosophy using New Relic for deep APM and system insights. Partner this with IncidentIO to ensure we catch and resolve issues before they impact customers.
Apply directly at LeanData →Create a free account for alerts like thisView LeanData immigration profile
This listing is sourced directly from LeanData's careers page and normalized into a canonical job model.