Microsoft

Microsoft

Senior Site Reliability Engineer

Redmond, WA, US · Senior

Sponsorship not specifiedDetected 11 hours ago
Site Reliability EngineeringNetwork EngineeringCommunicationProblem Solving

About the role

  • Enhancing customer facing experience by proactive alerting based on utilization, trends, resource health, etc.
  • 2+ years of operational experience in improving Service Reliability, Availability and Performance.
  • Understanding of Observability and MELT implementation patterns for large-scale services.

Responsibilities

  • Collaborating closely with engineering teams on building and enhancing tooling and automation solutions for faster resolution of issues impacting SLO's and averting incidents altogether when possible.
  • Ability to design and implement any changes to service telemetry for the automation to consume if it is not already available.
  • Analyze data and provide operational insights into customer experience to design and product teams, so that we can design features with supportability in mind.

Requirements

  • These requirements include, but are not limited to the following specialized security screenings: 4+ years of experience running large scale cloud services.

Company info

  • Collaborating with the customers to understand their pain points around supportability and SLO attainment and formulate strategies for addressing recurring issues in a sustainable way.
  • Communicate on a deeply technical level and be the single point of contact for interfacing with enterprise customers for handling service escalations and driving the issues to resolution.
  • Embody our culture and values.

This listing is sourced directly from Microsoft's careers page and normalized into a canonical job model.