Employerdirecthealthcare

Employerdirecthealthcare

Senior Site Reliability Engineer

Dallas, TX - Hybrid (3x in office/week) · Senior

Sponsorship not specifiedDetected 8 days ago
PythonBashPowerShellGitAWSGCPAzureKubernetesTerraformCI/CDGitHub ActionsPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringPlatform EngineeringIncident ResponseComplianceRecruitingHIPAACommunicationCollaboration

About the role

  • Lantern also pairs members with a dedicated care team, including Care Advocates and nurses, for the entirety of their care journey, helping them get back to good health, back to their families and back to work.
  • With convenient access to specialists nationwide, Lantern means quality care is within driving distance for most.
  • About You: - You use LOGIC in your decision making and understand that progress is critical to making change.

Responsibilities

  • Build and maintain observability platforms (monitoring, logging, alerting, tracing) using Datadog and Azure Monitor
  • Lead incident management processes using Rootly, including on-call rotations, runbooks, and post-incident reviews
  • Design and implement disaster recovery and business continuity strategies
  • Collaborate with development teams to improve service reliability through architecture reviews and chaos engineering
  • Optimize system performance, capacity planning, and cost efficiency for Azure infrastructure
  • Maintain and improve CI/CD pipelines to support safe, rapid deployments
  • By curating a Network of Excellence comprised of the nation's top specialists for surgery, cancer care, infusions and more, Lantern delivers excellent care with significant cost savings to employers and their workforces.

Requirements

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
  • 4+ years in SRE, DevOps, or production operations roles
  • Strong experience with observability tools (Datadog, Azure Monitor, Prometheus, Grafana, or similar)
  • Experience defining and managing SLOs/SLIs and error budgets
  • Proven incident management and on-call experience (Rootly or similar incident management platforms)
  • Deep experience with chaos engineering and reliability testing
  • Experience with Azure Kubernetes Service and containerized workloads

Nice to have

  • 3+ years with Microsoft Azure (AWS/GCP a plus)

Benefits

  • Medical Insurance
  • Dental Insurance
  • Vision Insurance
  • Short & Long Term Disability
  • Life Insurance
  • Flexible Time Off
  • Paid Parental Leave
  • Define and track SLOs/SLIs/error budgets for critical healthcare services

This listing is sourced directly from Employerdirecthealthcare's careers page and normalized into a canonical job model.