Zayo

Zayo

Principal Site Reliability Engineer

United States · Principal

Sponsorship not specified$117k-$180kDetected 22 days ago
PythonAWSCloud PlatformsDockerKubernetesAnsibleLinuxPrometheusGrafanaSite Reliability EngineeringTCP/IPDNSCiscoBGP/OSPFNetwork MonitoringLeadershipCommunicationCollaborationProblem SolvingCritical Thinking

About the role

  • Zayo's 141,000-mile network in North America and Europe includes extensive metro connectivity to thousands of buildings and data centers.
  • Zayo's communications infrastructure solutions include dark fiber, private data networks, wavelengths, Ethernet, and dedicated Internet access.
  • Zayo serves wireless and wireline carriers, media, tech, content, finance, healthcare and other large enterprises.

Responsibilities

  • Automation: Develop and implement automation solutions to streamline operations and reduce manual effort.
  • Monitoring and Alerting: Design and implement effective monitoring and alerting systems to proactively identify and address issues.
  • Incident Management: Own the incident lifecycle, from leading root cause analysis and resolution to implementing preventative measures to avoid future occurrences. Be on-call to diagnose and resolve critical service outages.
  • Scalability and Performance: Design and implement solutions to ensure our infrastructure can handle ever-growing demands while maintaining optimal application performance.
  • Develop and implement automation solutions to streamline operations and reduce manual effort.
  • Design and implement effective monitoring and alerting systems to proactively identify and address issues.
  • Own the incident lifecycle, from leading root cause analysis and resolution to implementing preventative measures to avoid future occurrences.
  • Design and implement solutions to ensure our infrastructure can handle ever-growing demands while maintaining optimal application performance.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience).
  • Minimum of ten (10) years of experience in a Site Reliability Engineering or related role.
  • Strong understanding of system administration, Linux, and scripting languages (Python and various shells).
  • Experience with monitoring platforms such as SevOne, Assure1, and Nagios and various vendor NMS systems.
  • Experience with a variety of cloud platforms and tools (AWS, Google, etc).
  • Experience with a variety of monitoring and alerting tools (Prometheus, Grafana, Cacti, etc.)
  • Strong working knowledge of networking concepts and application protocols, especially TCP/IP, BGP, DNS, TLS, and HTTP/S.
  • Experience with infrastructure management tools such as Ansible, Terrafor, Puppet, to deploy and manage infrastructure at scale.
  • Proven leadership skills, with the ability to mentor and inspire others.
  • Expert at developing automation tools for monitoring, alerting, and deployment to ensure efficient and reliable operations.
  • Expert at designing and implementing monitoring systems at scale.
  • Expert at container orchestration (Kubernetes and Docker).
  • Previous work in large scale distributed production environments.
  • Excellent problem-solving, analytical, and critical thinking skills.

Nice to have

  • Experience working with various vendor APIs (or netconf) including Nokia, Juniper, Fujitsu, Infinera, Cisco, and Ciena.
  • Experience with various network orchestration platforms such as Ciena Blue Planet MDSO, Cisco NSO, Nokia NSP, or others.

Compensation

  • $116,700 to $179,500 USD/annually.
  • The base pay range shown is a guideline and reasonable estimate for this role.
  • It takes into account the wide variety of factors that are considered in making compensation decisions.
  • Actual compensation offered may vary from the posted range based upon geographic location, work experience, skill level, certifications, and other business and organizational needs.
  • Non- sales roles may be eligible to participate in a discretionary annual incentive plan.
  • Sales roles may be eligible to participate in a sales incentive plan.

Benefits

  • Excellent Health, Dental & Vision Insurance
  • Retirement 401(k) Savings Plan
  • Generous paid time off policy including paid parental leave
  • Additionally, this position may be eligible for certain benefits, such as health insurance, life insurance, disability retirement plans, paid time off.
  • Benefits, Rewards & Wellness

Company info

  • Company Description
  • Zayo provides mission-critical bandwidth to the world's most impactful companies, fueling the innovations that are transforming our society.
  • Zayo is seeking a talented Principal Site Reliability Engineer to play a critical role in ensuring the uptime, performance, and scalability of our critical infrastructure.
  • Do you dream in high scalable systems, thrive in fast-paced environments and enjoy tackling complex technical challenges?
  • Are you passionate about building and maintaining highly reliable, scalable systems?
  • If so, then join our team as a Principal Site Reliability Engineer (SRE)!

Equal opportunity

  • Zayo is an equal opportunity employer to all protected groups, including protected veterans and individuals with disabilities.
  • If you are an individual with a disability and would like to request a reasonable accommodation as part of the employment selection process, please contact us.

This listing is sourced directly from Zayo's careers page and normalized into a canonical job model.