O.C. Tanner
Manager, Site Reliability Engineering
USA - Utah-Salt Lake City-Headquarters
Sponsorship not specifiedDetected 1 day ago
PythonGoPostgreSQLRedisElasticsearchAWSKubernetesTerraformDatadogDevOpsSite Reliability EngineeringPlatform EngineeringKafkaIncident ResponseBudgetingPerformance ManagementPlaywrightLeadership
About the role
- O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences.
- Join us as we help people all over the world thrive at work.
Responsibilities
- Lead, mentor, and develop a team of Site Reliability Engineers, fostering a culture of reliability, accountability, operational excellence, and continuous improvement.
- Partner with Engineering, Product, and Support leaders to drive shared ownership of production services and embed reliability, observability, and operational excellence throughout the software development lifecycle.
- Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar tools, establishing enterprise standards for metrics, logs, traces, alerting, and Service Level Objectives (SLOs).
- Manage team capacity, hiring, performance management, career development, budgeting, and workforce planning to ensure effective support of business-critical services.
- Demonstrated ability to lead cross-functional initiatives and influence stakeholders across Engineering, Product, and Support organizations.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines, including 2+ years in a technical leadership or people management role.
- Proven experience leading teams responsible for production operations, reliability engineering, incident management, and operational excellence.
- Experience designing and implementing SRE practices, reliability programs, or operational maturity initiatives within growing engineering organizations.
- Experience operating large-scale, customer-facing SaaS platforms with high availability, performance, and scalability requirements.
- Strong understanding of modern software engineering practices and partnering with development teams to build reliable, resilient systems.
- Hands-on experience with observability platforms such as OpenTelemetry, Datadog, Coralogix, or similar technologies.
- Strong knowledge of AWS and Kubernetes in production environments.
Nice to have
- Experience leading distributed or globally dispersed engineering teams.
- Experience with multiple cloud providers or cloud-agnostic platform architectures.
- Familiarity with security, compliance, governance, and operational risk management frameworks.
- Proficiency with modern Infrastructure-as-Code and technologies such as Terraform, Golang, Python, Playwright, and Performance Monitoring tools.
- Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
- Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.
- Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics.
Benefits
- Establish team priorities, goals, and success metrics aligned with business objectives, customer needs, and platform health.
- Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective operational handoffs, and consistent service ownership.
- Own on-call programs, incident management practices, and operational health metrics, driving improvements in alert quality, operational efficiency, and toil reduction.
Apply directly at O.C. Tanner →Create a free account for alerts like thisView O.C. Tanner immigration profile
This listing is sourced directly from O.C. Tanner's careers page and normalized into a canonical job model.