Microsoft
Senior Site Reliability Engineer
United States, Washington, Redmond · Senior
Sponsorship not specified$120k-$235kDetected 11 hours ago
SQLPostgreSQLAzureSite Reliability EngineeringRESTData AnalysisData EngineeringNetwork EngineeringCommunicationProblem Solving
About the role
- Microsoft's Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence.
- The products our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI.
Responsibilities
- Collaborating closely with engineering teams on building and enhancing tooling and automation solutions for faster resolution of issues impacting SLO's and averting incidents altogether when possible.
- Ability to design and implement any changes to service telemetry for the automation to consume if it is not already available.
- Analyze data and provide operational insights into customer experience to design and product teams, so that we can design features with supportability in mind.
- Microsoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further.
- We store and manage data in a structured way to enable multitude of applications across various industries.
- It is designed to enable developers to build planet-scale applications.
- As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals.
- Our mission is to build the data platform for the age of AI, powering a new class of data-first applications and driving a data culture.
- Within Azure Data, the databases team builds and maintains Microsoft's operational Database systems.
Requirements
- OR Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration.
- Ability to meet Microsoft, customer and/or government security screening requirements are required for this role.
- Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
- Required/Minimum Qualifications:
- 6+ years technical experience in software engineering, network engineering, or systems administration
- OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration
- This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
Nice to have
- 4+ years of experience running large scale cloud services.
- 2+ years of operational experience in improving Service Reliability, Availability and Performance.
- Understanding of Observability and MELT implementation patterns for large-scale services.
- Experience in Logic Apps and authoring Jupyter Notebooks.
- Experience in analyzing, troubleshooting, and automating root cause analysis and mitigation of incidents impacting large-scale distributed systems.
- Systematic problem-solving approach, coupled with effective communication skills and a sense of curiosity.
- Ability to deal with the ambiguity associated with working in a fast-paced environment.
- Influencing the product architecture and roadmap to make sure the customer-experienced supportability is always a key consideration when evolving the product.
Skills
- This is a world of more possibilities, more innovation, more openness, and the sky is the limit thinking in a cloud-enabled world.
- Azure Cosmos DB is Microsoft's next generation of globally distributed, massively scalable, multi-model cloud database service.
- Cosmos DB is a database of choice for the spectrum spanning from the hobbyist developer to the largest of Fortune 500 companies.
Compensation
- $120k-$235k
Benefits
- Enhancing customer facing experience by proactive alerting based on utilization, trends, resource health, etc.
Company info
- We are looking for a self-driven Senior Site Reliability Engineer (SRE) who likes taking a data driven and systems-based approach to solve Service Reliability problems.
- Collaborating with the customers to understand their pain points around supportability and SLO attainment and formulate strategies for addressing recurring issues in a sustainable way.
- Communicate on a deeply technical level and be the single point of contact for interfacing with enterprise customers for handling service escalations and driving the issues to resolution.
- Embody our culture and values.
- Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
- We are on a journey to enable developer friendly, mission-critical, AI enabled operational databases across relational, non-relational and OSS offerings.
Visa & Work Authorization
- All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status
Apply directly at Microsoft →Create a free account for alerts like thisView Microsoft immigration profile
This listing is sourced directly from Microsoft's careers page and normalized into a canonical job model.