Kontakt.io

Kontakt.io

SRE Leader

New York

Sponsorship not specifiedDetected 9 days ago
Distributed SystemsAWSCloud PlatformsDockerKubernetesTerraformCI/CDPrometheusGrafanaSite Reliability EngineeringCybersecurityNetwork SecurityIncident ResponseComplianceHIPAAHL7/FHIREHR/EMRPatient CareLeadershipMentoring

About the role

  • About Kontakt.io http://Kontakt.io Every day, intelligent software orchestrates the physical world around us - from matching drivers with riders to optimizing global supply chains.
  • Yet inside hospitals, where every second and every decision can affect a patient's life, operations are still spread across dozens of disconnected systems.
  • At Kontakt.io http://Kontakt.io, we're changing that.

Responsibilities

  • Design and implement self-healing, fault-tolerant systems to prevent failures before they happen.
  • Architect and manage scalable cloud infrastructure (AWS) for massive real-time data processing.
  • Optimize containerized environments (Kubernetes, Docker) to support multi-region deployments.
  • Lead the adoption of infrastructure as code (Terraform) to fully automate infrastructure management.
  • Build and refine a world-class monitoring, alerting, and logging system using Prometheus, Grafana, OpenTelemetry, and Datadog.
  • Lead incident response and on-call operations, reducing mean time to detection (MTTD) and mean time to resolution (MTTR).
  • Partner with Security & Compliance teams to ensure infrastructure meets HIPAA and SOC 2 standards.
  • Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering.
  • Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.
  • That intelligence powers the execution layer hospitals have been missing - helping care teams make smarter decisions and deliver better patient care.

Requirements

  • 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure.
  • Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems.
  • Hands-on experience with incident management, postmortems, and building resilient systems.
  • Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.).
  • Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC 2).
  • Proven success scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare.
  • Strong background in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools.
  • A mature leadership approach, with the ability to drive technical strategy while growing and mentoring a high-performance SRE team.
  • Bonus Points If You Have:
  • Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability.
  • Own Mission-Critical Reliability - Ensure hospitals and care facilities always stay online with a 99.99% uptime healthcare platform.

Nice to have

  • Expertise in real-time distributed systems, event-driven architectures, or large-scale data pipelines.
  • Prior experience leading on-call rotations and major incident management processes.
  • Scale AI-Powered Infrastructure - Work on real-time automation and self-healing cloud systems that orchestrate care delivery.
  • Automation-First Culture - Minimize manual ops with cutting-edge automation, observability, and incident response strategies.

Compensation

  • We've more than doubled our revenue over the past year and are on track to surpass $70M in annual recurring revenue - not because we're following a market, but because we're defining one.

Benefits

  • You will lead and scale the SRE team to ensure our infrastructure stays ahead of demand, operates efficiently, and meets the needs of our growing healthcare customers.
  • Ensure 99.99% uptime across our cloud platform, meeting strict SLAs for healthcare customers.
  • Lead disaster recovery and business continuity planning to ensure critical healthcare services are always available.

This listing is sourced directly from Kontakt.io's careers page and normalized into a canonical job model.