Kong

Kong

Senior Site Reliability Engineer, Kong Konnect

Canada · Senior

Sponsorship not specifiedDetected 258 days ago
PythonGoBashPostgreSQLRedisAWSGCPAzureKubernetesTerraformHelmCI/CDLinuxPrometheusGrafanaDatadogSite Reliability EngineeringKafkaIncident ResponseDNSFirewall

About the role

  • You'll work on everything from multi-region Kubernetes clusters to service mesh and gateway architectures, ensuring the reliability, scalability, and security of Kong's SaaS offerings.
  • This is a hands-on role ideal for engineers who thrive on running production SaaS systems at scale, automating operations, and continuously improving performance, resilience, and deployment pipelines.
  • Operate and scale Kong's global SaaS platform (Konnect), ensuring reliability, availability, and performance across regions and clouds.

Responsibilities

  • Build, automate, and maintain Kubernetes-based infrastructure and deployment workflows using Terraform/Terragrunt, Helm, and ArgoCD.
  • Design, maintain, and optimize multi-region data and caching layers - including PostgreSQL, Redis, ClickHouse, and Druid - for high availability and low latency.
  • Develop and maintain CI/CD pipelines and GitOps workflows to automate service delivery and ensure consistent infrastructure changes.
  • Collaborate closely with development and security teams to ensure smooth operation of SaaS services in compliance with reliability, security, and regulatory standards.
  • Participate in a global 24/7 on-call rotation and drive continuous improvement of operational playbooks and postmortem practices.
  • Lead and contribute to scaling initiatives that improve elasticity, reliability, and cost-efficiency across the SaaS platform.

Requirements

  • Proven experience managing SaaS or PaaS systems at enterprise scale (multi-region, multi-tenant, secure environments).
  • Deep expertise in Kubernetes, including debugging cluster/networking issues and designing for fault tolerance and scalability.
  • Strong proficiency with Infrastructure as Code tools like Terraform or Terragrunt.
  • Experience with CI/CD pipelines and GitOps workflows (ArgoCD, Atlantis, Helm).
  • Proficiency in one or more programming languages (Go, Python, Bash) for automation and tooling.
  • Familiarity with streaming systems like Kafka and observability platforms (Datadog, Prometheus, Grafana).

Nice to have

  • Hands-on experience with Kong Gateway, Kong Mesh, or similar service connectivity technologies.
  • Experience managing PostgreSQL and Redis in multi-region configurations.
  • Working knowledge of AWS networking (PrivateLink, Transit Gateway, VPC Peering, Firewalls), Azure VNet, or GCP NCC.
  • Strong understanding of disaster recovery, resiliency testing, and compliance-driven reliability practices.

Skills

  • For more information, visit www.konghq.com http://www.konghq.com.
  • Experiencing working with API gateway and service mesh technologies
  • Experience operating ClickHouse, Druid, or other time-series and analytics databases.

Company info

  • Nobody checks every box - we're looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.

This listing is sourced directly from Kong's careers page and normalized into a canonical job model.