Kody
Senior Site Reliability Engineer- San Francisco, CA, the US
San Francisco, California, United States · Senior
Sponsorship not specifiedDetected 29 days ago
Distributed SystemsPostgreSQLRedisAWSCloud PlatformsKubernetesTerraformLinuxDevOpsSite Reliability EngineeringPlatform EngineeringKafkaIncident ResponseComplianceLeadershipMentoring
About the role
- Participate in a follow-the-sun production on-call rotation as a primary incident responder.
- Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
- Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization.
Requirements
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems.
- Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms.
- Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements.
- Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence.
- Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.
This listing is sourced directly from Kody's careers page and normalized into a canonical job model.