Kody

Kody

Senior Site Reliability Engineer- San Francisco, CA, the US

San Francisco, California, United States · Senior

Sponsorship not specifiedDetected 29 days ago
Distributed SystemsPostgreSQLRedisAWSCloud PlatformsKubernetesTerraformLinuxDevOpsSite Reliability EngineeringPlatform EngineeringKafkaIncident ResponseComplianceLeadershipMentoring

About the role

  • Participate in a follow-the-sun production on-call rotation as a primary incident responder.
  • Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
  • Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization.

Requirements

  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems.
  • Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms.
  • Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements.
  • Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence.
  • Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.

This listing is sourced directly from Kody's careers page and normalized into a canonical job model.