Microsoft

Microsoft

Posted via Microsoft Careers

Sr. Incident Management Engineer

Posted Aug 11, 2026

Role at a glance

Salary
$119.8K – $234.7K/yr
Location
Redmond, Washington, United States
Work arrangement
On-site
Employment
Full-time
Education
Bachelor's, Master's

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

The Senior Incident Management Engineer leads resolution of critical service and infrastructure incidents across Azure and cloud platforms. The role partners with engineering teams, service owners, SRE teams, and datacenter operations to improve reliability, operational maturity, service resilience, and customer impact during incidents.

What You'll Do

  • Lead mitigation and resolution of high-severity service and infrastructure incidents.
  • Coordinate cross-functional engineering teams during live-site events and drive communication, decision-making, escalation, and mitigation.
  • Lead post-incident reviews and root cause investigations, and drive corrective and preventive actions.
  • Improve incident management processes, response playbooks, escalation paths, and incident response effectiveness.
  • Track and analyze operational metrics and trends, and identify systemic reliability risks.
  • Support operational readiness reviews and outage preparedness activities.

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

Master's degree in Electrical Engineering, Computer Engineering, or a related field with 3+ years of technical engineering experience, or bachelor's degree in one of those fields with 5+ years of technical engineering experience, or equivalent experience.

Required

  • Master's Degree in Electrical Engineering, Computer Engineering, or related field
  • Bachelor's Degree in Electrical Engineering, Computer Engineering, or related field
  • 3+ years technical engineering experience with a master's degree
  • 5+ years technical engineering experience with a bachelor's degree
  • Ability to meet Microsoft, customer and/or government security screening requirements

Preferred

  • Familiarity with HW (EE, thermal, mechanical) issues
  • Experience in incident management, cloud operations, Site Reliability Engineering, infrastructure engineering, or related disciplines
  • Experience leading complex technical incidents involving multiple teams
  • Proven troubleshooting and analytical skills
  • Proven communication and stakeholder management capabilities
  • Experience with Microsoft IcM or other enterprise incident management systems
  • Familiarity with KQL, Azure Monitor, Power BI, or operational analytics tools
  • Experience automating operational processes using scripting or automation frameworks

Original job description

Content provided by the employer

Overview

Microsoft is seeking a Senior Incident Management (IcM) Engineer to lead the resolution of critical service and infrastructure incidents across Azure and cloud platforms. This role combines incident leadership, technical problem-solving, operational excellence, and reliability improvement. 

The ideal candidate has experience leading high-severity incidents, driving root cause analysis, influencing engineering teams, and improving the operational maturity of large-scale services. 

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. 



Responsibilities
 

Incident Leadership:

  • Lead mitigation and resolution of high-severity service and infrastructure incidents. 

  • Coordinate cross-functional engineering teams during live-site events. 

  • Drive clear communication and decision-making during incident response. 

  • Ensure timely escalation, mitigation, and customer impact reduction. 

Root Cause & Reliability Improvement:

  • Lead post-incident reviews and root cause investigations. 

  • Drive corrective and preventive actions to reduce recurring incidents. 

  • Identify systemic reliability risks and recommend improvements. 

  • Partner with engineering teams to improve service resilience. 

Operational Excellence:

  • Improve incident management processes, response playbooks, and escalation paths. 

  • Track and analyze operational metrics and trends. 

  • Drive improvements in incident response effectiveness and service availability. 

  • Support operational readiness reviews and outage preparedness activities. 

Leadership & Collaboration: 

  • Influence engineering teams using data and operational insights. 

  • Partner with service owners, SRE teams, datacenter operations, and engineering organizations. 

  • Foster a culture of accountability, reliability, and customer focus. 



Qualifications

Required Qualifications:

  • Master's Degree in Electrical Engineering, Computer Engineering, or related field AND 3+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, or related field AND 5+ years technical engineering experience
    • OR equivalent experience.  

 

Other Requirements:
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:
 
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter

 

Preferred Qualifications:

  • Familiarity with HW (EE, thermal, mechanical) issues and high-level ability to identify and route to SME as appropriate  

  • Experience in incident management, cloud operations, Site Reliability Engineering, infrastructure engineering, or related disciplines. 

  • Experience leading complex technical incidents involving multiple teams. 

  • Proven troubleshooting and analytical skills.

  • Proven communication and stakeholder management capabilities. 

  • Experience with Microsoft IcM or other enterprise incident management systems. 

  • Familiarity with KQL, Azure Monitor, Power BI, or operational analytics tools. 

  • Experience automating operational processes using scripting or automation frameworks. 

What Success Looks Like 

  • Independently leads complex, high-impact incidents. 

  • Drives measurable reductions in recurring operational issues. 

  • Influences engineering teams to improve service reliability. 

  • Uses operational insights to improve customer experience and platform stability. 

  • Recognized as a trusted leader during critical incidents and escalations. 

Target Candidate: A senior technical leader who thrives in high-pressure situations, can coordinate across organizations, and is passionate about improving the reliability and operational excellence of Microsoft's cloud services. 



Electrical Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Microsoft

About the company

Microsoft

Large Enterprise

Microsoft is a global technology leader that empowers individuals and organizations to achieve more through innovative software, services, and devices. Founded in 1975, the company is best known for its flagship products like the Windows operating system and Microsoft Office suite. In addition to personal computing, Microsoft is a leader in cloud computing with its Azure platform, providing a range of solutions for businesses to enhance productivity and efficiency. With a strong commitment to sustainability and accessibility, Microsoft continues to drive technological advancements that shape the future of work and learning.