Ncr
Senior Site Reliability Engineer – Unified Observability
ATLANTA, GA, USA · Senior
Sponsorship not specifiedDetected 1 day ago
PythonGoPowerShellGCPAzureCloud PlatformsKubernetesTerraformCI/CDPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringPlatform EngineeringStakeholder ManagementRecruitingLeadershipCommunication
About the role
- This role will serve as a key technical leader responsible for defining standards, influencing architecture decisions, and enabling proactive operations through unified monitoring, telemetry, automation, and AI-driven insights.
- Offers of employment are conditional upon passage of screening criteria applicable to the job EEO Statement Integrated into our shared values is NCR Voyix's commitment to equal employment opportunity.
Responsibilities
- Establish and drive enterprise observability standards for monitoring, logging, distributed tracing, telemetry, and operational analytics.
- Partner cross-functionally with Engineering, Infrastructure, Security, Product, and Operations teams to identify reliability risks and drive operational excellence initiatives.
- Lead efforts to improve incident prevention, detection, response, and recovery through intelligent alerting, automation, event correlation, and observability best practices.
- Support and drive enterprise initiatives involving AI-driven observability, predictive analytics, anomaly detection, and event intelligence.
Requirements
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
- 10+ years of experience in Site Reliability Engineering, Cloud Engineering, Platform Engineering, DevOps, or related technical disciplines.
- Deep expertise with Kubernetes platforms, including AKS and GKE.
- Strong experience with Azure and Google Cloud Platform services and architectures.
- Hands-on experience with enterprise observability tools such as Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic, or similar platforms.
- Advanced knowledge of monitoring, logging, telemetry collection, distributed tracing, and observability engineering principles.
- Strong automation and Infrastructure as Code expertise using Terraform and related tools.
- Proficiency developing automation solutions using Python, Go, PowerShell, or similar languages.
- Experience integrating observability solutions into CI/CD pipelines and modern DevOps workflows.
- Strong communication, leadership, and stakeholder management skills with the ability to translate technical concepts for both technical and business audiences.
Nice to have
- Experience leading enterprise observability transformations or platform modernization initiatives.
- Experience with AI Ops, event correlation, operational analytics, and predictive monitoring capabilities.
- Knowledge of ServiceNow integrations and ITSM/ITOM processes.
- Experience supporting highly available, customer-facing SaaS platforms at scale.
- One or more cloud certifications (Azure, Google Cloud, Kubernetes, or related technologies).
- Offers of employment are conditional upon passage of screening criteria applicable to the job
- Integrated into our shared values is NCR Voyix's commitment to equal employment opportunity.
- NCR Voyix is committed to being a globally inclusive company where all people are treated fairly, recognized for their individuality, promoted based on performance and encouraged to strive to reach their full potential.
Benefits
- Develop and maintain executive, operational, and engineering dashboards that provide real-time visibility into infrastructure, applications, platform health, customer experience, and business transactions.
- Define, evangelize, and implement reliability frameworks including SLIs, SLOs, error budgets, operational KPIs, and service health metrics.
- Experience defining and operationalizing SLIs, SLOs, error budgets, reliability metrics, and service health frameworks.
This listing is sourced directly from Ncr's careers page and normalized into a canonical job model.