RaceTrac

RaceTrac

Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert

200 Galleria Parkway SE Suite 900 Atlanta, GA 30339 · Senior

Sponsorship not specifiedDetected 6 days ago
PythonJavaC++iOSAndroidDistributed SystemsSQLDatabricksAzureTerraformAnsibleCI/CDJenkinsDevOpsSite Reliability EngineeringKQLIncident ResponseStakeholder ManagementLoad BalancingCommunicationCollaborationProblem Solving

About the role

  • Fueled by Growth, Driven by You At RaceTrac, our people make the difference.
  • With four operating divisions RaceTrac, RaceWay, Energy Dispatch, and Gulf - there's always a new challenge to take on and a new path to pursue.
  • Join us and discover how far your career can go.

Responsibilities

  • Design, develop, and optimize enterprise observability solutions using Dynatrace.
  • Develop advanced Dynatrace DQL queries, dashboards, workflows, alerts, and analytics.
  • Implement intelligent monitoring strategies for applications, APIs, integrations, Azure services, mobile platforms, and distributed systems.
  • Build and maintain monitoring solutions using:
  • Engages in and improve the whole lifecycle of services-from inception and design, deployment, operation, and refinement.
  • Develops software and provide hands-on technical knowledge to design, deploy, and optimize large-scale, massively distributed, fault-tolerant systems.
  • Supports services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity planning, automation, pipelining and launch reviews.
  • Improves, tunes and performs operational efficiency within the Windows based infrastructure and production environment.
  • Collaborates with development teams to support the current environment as we transform into a cloud architecture and provides resources "as a service" to developers.
  • Develop proactive monitoring for cloud services, integrations, APIs, and backend processing systems.

Requirements

  • Experience programming in at least one of the following languages: C, C++, Java, Python, or Go.
  • Minimum 4 years of working experience in Azure.
  • Experience with Jenkins or similar application.
  • General knowledge of Infrastructure as Code tools and Config management tools such as (Terraform/Ansible/Chef/Puppet/SCCM).
  • Comfort with large-scale production systems and technologies (load balancing, monitoring, distributed system and configuration management.
  • Expertise in designing, analyzing, and troubleshooting.
  • 10+ years of overall IT experience.
  • Expert-level hands-on experience with Dynatrace.
  • Advanced expertise in Dynatrace Query Language (DQL).
  • Strong hands-on expertise in Azure Kusto Query Language (KQL).
  • Bachelor's degree from an accredited college or university in Computer Science or related field preferred. Equivalent practical experience will be considered.
  • Minimum 4 years of working experience in Azure. Experience with Jenkins or similar application.
  • Comfort with large-scale production systems and technologies (load balancing, monitoring, distributed system and configuration management. Expertise in designing, analyzing, and troubleshooting.
  • Ability to debug, optimize code, and automate routine tasks.
  • Systematic problem-solving approach, coupled with effective communication skills and a sense of drive.
  • Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity management and launch reviews.
  • Demonstrated history of living the values that are important to RaceTrac: Honesty, Efficiency, Attitude, Respect, Teamwork.
  • All qualified applicants will receive consideration for employment with RaceTrac without regard to their race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.

Nice to have

  • Azure API Management (APIM)
  • Strong understanding of API architectures, API Gateways, and backend integrations.
  • Strong ability to read, analyze, and understand.NET application code.
  • Experience troubleshooting and deploying Azure Functions and cloud-native applications.
  • Experience enabling observability and telemetry for mobile applications on iOS and Android.
  • Understanding of mobile telemetry, crash analytics, API monitoring, and end-user experience monitoring.
  • Strong understanding of distributed systems and enterprise application architectures.
  • Experience with OpenTelemetry implementation and instrumentation.

Skills

  • Azure Monitor
  • Application Insights
  • Azure Log Analytics
  • Azure KQL
  • Monitor and troubleshoot Azure Function Apps, App Services, APIs, integrations, and backend services.
  • Analyze telemetry, traces, logs, metrics, and distributed transactions to identify root causes and performance bottlenecks.
  • Troubleshoot cloud-native applications and Azure infrastructure issues.
  • Monitor and troubleshoot Azure API Management (APIM), API Gateways, API endpoints, and integrations.
  • Understand end-to-end API transaction flows and dependency mapping.
  • Diagnose latency issues, transaction failures, authentication issues, and backend service degradation.
  • Enable telemetry, monitoring, tracing, and performance analysis for iOS and Android applications.
  • Analyze mobile-to-backend transaction flows and end-user experience metrics.

Benefits

  • Maintains services once they are live by measuring/monitoring availability, latency, and overall system health.
  • Reduces manual intervention and turn-around time to solve for repetitive problems while automating and monitoring the health of our sites and services.

Company info

  • RaceTrac Company Overview

Equal opportunity

  • Honesty, Efficiency, Attitude, Respect, Teamwork.

This listing is sourced directly from RaceTrac's careers page and normalized into a canonical job model.