Microsoft

Microsoft

Posted via Microsoft Careers

Software Engineer II

Posted Sep 15, 2026

Role at a glance

Job function
Software Engineering & IT Software Engineering
Salary
$102.1K – $202.2K/yr
Location
Redmond, Washington, United States
Work arrangement
On-site
Employment
Full-time
Education
Bachelor's

Spotted an issue?

We’ll check it against the original posting.

Log in to report

About the role

Original posting provided by Microsoft

View original
Overview
The AI Infrastructure team is responsible for building and operating the large-scale, reliable, and efficient GPU-based clustering infrastructure that powers Microsoft’s AI/ML ecosystem. We host the training and inference platforms behind many of Microsoft’s flagship AI offerings, including Azure OpenAI Service, M365 Copilot & Copilot Tuning, GitHub Copilot, Azure AI Foundry’s inference and fine-tuning services for both OpenAI and open-source models; as well as the mature Azure ML Services, which provide data scientists and developers a rich experience for defining, training, fine-tuning, deploying, monitoring, and consuming machine learning models. Our infrastructure enables AI innovation at hyperscale and supports some of the most demanding workloads and business groups across the company. We engage directly with some of the major internal research and applied AI/ML groups using these services, including Microsoft Research, M365, Microsoft Security, and the Bing WebXT team.
 
The AI Infra team is looking for a talented Software Engineer II, with initial focus on the Scheduler subsystem. The scheduler is the “brains” of the AI Infra control plane. It governs access to the GPU, NPU and CPU capacity of the platform according to a complex system of workload preference rules, placement constraints, optimization objectives, and dynamically interacting policies aimed to maximize hardware utilization and fulfill greatly varying needs of users and the AI platform partner services in terms of workload types, prioritization, and capacity targeting flexibility. The scheduler’s set of capabilities is broad and ambitions. It manages quota, capacity reservations, SLA tiers, preemption, auto-scaling, and a wide range of configurable policies. It is both a workload-aware and topology-aware scheduler (down to the level of cluster racks and nodes). Global scheduling is a distinctive major feature that overcomes the regional segmentation of the Azure compute fleet by treating the GPU capacity as a single global virtual pool, which greatly increases capacity availability and utilization for major classes of AI/ML workload. We have achieved this capability by avoiding a significant global single point of failure, based on regional instances of the scheduler service interacting via peer-to-peer protocols for sharing capacity inventory and coordinating handoff of jobs for scheduling. Our system manages significant amount of GPU capacity even outside Azure datacenters, through a unified model and operational process and highly generalized, flexible workload scheduling capabilities.
 
To be able to manage the inherent complexity of the Scheduler subsystem and enable it to meet the stringent expectations of high service reliability, availability, and throughput, we emphasize rigorous engineering, utmost precision and quality, and strong ownership—from feature design to livesite. Quality mindset, attention to detail, development process rigor, and data-driven design and problem-solving skills are key for success in our mission-critical control plane space. We enjoy great creative freedom and thrive on the capacity management and workload scheduling challenges posed to us by the various partner teams.
 

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.



Responsibilities
  • Work on the design and development of the core AI Infrastructure distributed and in-cluster services that support large scale AI training and inferencing.
  • Develop, test, and maintain control plane services written in C#, hosted on Service Fabric or Kubernetes (AKS) clusters.
  • Enhance systems and applications to ensure high stability, efficiency and maintainability, low latency, tight cloud security.
  • Provide operational support and DRI (on-call) responsibilities for the service.
  • Develop and foster a deep understanding of the AI/ML concepts, use cases, and relevant services used by our customers. Be an AI-first developer, making productive use of the available tools and actively engaging in experimentation and learning.
  • Collaborate closely with service engineers, product managers, and internal applied research and data science teams within Microsoft to build better solutions together.
  • Provide vision, expertise, and technical leadership to other team members.
  • Embody our culture and values
 


Qualifications

Required Qualifications: 

  • Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience. 

Other Requirements:

Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:

Microsoft Cloud Background Check:

- This position will be required to pass the Microsoft background and Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Preferred Qualifications: 

  • Hands-on (devops) experience with larger-scale, high-availability cloud services at the PaaS or IaaS level, based on microservices architecture, ideally related to AI infrastructure or workload hosting
  • Proficiency with use of complex data structures and algorithms, preferably in the setting of a resource allocator/scheduler, workflow/execution orchestration engine, database engine, or similar
  • Proficiency and thoroughness in unit testing and testability techniques
  • Agentic development skills
  • Experience with building and operating “stateful” and critical control plane services; handling challenges with data size and data partitioning; advanced use of a NoSQL cloud database
  • Service reliability and fundamentals engineering; instrumentation for KPIs or performance analysis; demonstrated service and code quality mindset
  • Applied knowledge of Kubernetes: service model, workload packaging and deployment, programmatic extensibility (CRDs, operators); or equivalent knowledge of Service Fabric; experience with any service mesh
  • Data-driven design and troubleshooting and data analytics skills, ideally with Kusto
#AIINFRA
 


Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Microsoft

About the company

Microsoft

Large Enterprise

Microsoft is a global technology leader that empowers individuals and organizations to achieve more through innovative software, services, and devices. Founded in 1975, the company is best known for its flagship products like the Windows operating system and Microsoft Office suite. In addition to personal computing, Microsoft is a leader in cloud computing with its Azure platform, providing a range of solutions for businesses to enhance productivity and efficiency. With a strong commitment to sustainability and accessibility, Microsoft continues to drive technological advancements that shape the future of work and learning.