Zayo
Principal Site Reliability Engineer
United States · Principal
Sponsorship not specified$117k-$180kDetected 22 days ago
PythonAWSCloud PlatformsDockerKubernetesAnsibleLinuxPrometheusGrafanaSite Reliability EngineeringTCP/IPDNSCiscoBGP/OSPFNetwork MonitoringLeadershipCommunicationCollaborationProblem SolvingCritical Thinking
About the role
- Zayo's 141,000-mile network in North America and Europe includes extensive metro connectivity to thousands of buildings and data centers.
- Zayo's communications infrastructure solutions include dark fiber, private data networks, wavelengths, Ethernet, and dedicated Internet access.
- Zayo serves wireless and wireline carriers, media, tech, content, finance, healthcare and other large enterprises.
Responsibilities
- Automation: Develop and implement automation solutions to streamline operations and reduce manual effort.
- Monitoring and Alerting: Design and implement effective monitoring and alerting systems to proactively identify and address issues.
- Incident Management: Own the incident lifecycle, from leading root cause analysis and resolution to implementing preventative measures to avoid future occurrences. Be on-call to diagnose and resolve critical service outages.
- Scalability and Performance: Design and implement solutions to ensure our infrastructure can handle ever-growing demands while maintaining optimal application performance.
- Develop and implement automation solutions to streamline operations and reduce manual effort.
- Design and implement effective monitoring and alerting systems to proactively identify and address issues.
- Own the incident lifecycle, from leading root cause analysis and resolution to implementing preventative measures to avoid future occurrences.
- Design and implement solutions to ensure our infrastructure can handle ever-growing demands while maintaining optimal application performance.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience).
- Minimum of ten (10) years of experience in a Site Reliability Engineering or related role.
- Strong understanding of system administration, Linux, and scripting languages (Python and various shells).
- Experience with monitoring platforms such as SevOne, Assure1, and Nagios and various vendor NMS systems.
- Experience with a variety of cloud platforms and tools (AWS, Google, etc).
- Experience with a variety of monitoring and alerting tools (Prometheus, Grafana, Cacti, etc.)
- Strong working knowledge of networking concepts and application protocols, especially TCP/IP, BGP, DNS, TLS, and HTTP/S.
- Experience with infrastructure management tools such as Ansible, Terrafor, Puppet, to deploy and manage infrastructure at scale.
- Proven leadership skills, with the ability to mentor and inspire others.
- Expert at developing automation tools for monitoring, alerting, and deployment to ensure efficient and reliable operations.
- Expert at designing and implementing monitoring systems at scale.
- Expert at container orchestration (Kubernetes and Docker).
- Previous work in large scale distributed production environments.
- Excellent problem-solving, analytical, and critical thinking skills.
Nice to have
- Experience working with various vendor APIs (or netconf) including Nokia, Juniper, Fujitsu, Infinera, Cisco, and Ciena.
- Experience with various network orchestration platforms such as Ciena Blue Planet MDSO, Cisco NSO, Nokia NSP, or others.
Compensation
- $116,700 to $179,500 USD/annually.
- The base pay range shown is a guideline and reasonable estimate for this role.
- It takes into account the wide variety of factors that are considered in making compensation decisions.
- Actual compensation offered may vary from the posted range based upon geographic location, work experience, skill level, certifications, and other business and organizational needs.
- Non- sales roles may be eligible to participate in a discretionary annual incentive plan.
- Sales roles may be eligible to participate in a sales incentive plan.
Benefits
- Excellent Health, Dental & Vision Insurance
- Retirement 401(k) Savings Plan
- Generous paid time off policy including paid parental leave
- Additionally, this position may be eligible for certain benefits, such as health insurance, life insurance, disability retirement plans, paid time off.
- Benefits, Rewards & Wellness
Company info
- Company Description
- Zayo provides mission-critical bandwidth to the world's most impactful companies, fueling the innovations that are transforming our society.
- Zayo is seeking a talented Principal Site Reliability Engineer to play a critical role in ensuring the uptime, performance, and scalability of our critical infrastructure.
- Do you dream in high scalable systems, thrive in fast-paced environments and enjoy tackling complex technical challenges?
- Are you passionate about building and maintaining highly reliable, scalable systems?
- If so, then join our team as a Principal Site Reliability Engineer (SRE)!
Equal opportunity
- Zayo is an equal opportunity employer to all protected groups, including protected veterans and individuals with disabilities.
- If you are an individual with a disability and would like to request a reasonable accommodation as part of the employment selection process, please contact us.
This listing is sourced directly from Zayo's careers page and normalized into a canonical job model.