Full Job Description
We are looking for a Principal Software Engineer - HW/SW in Fleet Infrastructure to provide technical leadership for the evolution of Substrate's hardware and firmware platforms. In this role, you will shape New Product Introduction strategy, hardware and firmware health architecture, live-site reliability, and the future direction of AI-optimized infrastructure.
You will work across Microsoft 365, Azure infrastructure, hardware engineering teams, silicon providers, OEM(original equipment manufacturer)/ODM (original design manufacturer) partners, and operations organizations to define architecture, influence platform strategy, and solve complex infrastructure problems at global scale.
This role requires strong technical depth, architectural judgment, and the ability to create clarity across highly ambiguous, cross-organizational problems. Internal guidance for Principal engineers emphasizes broad scope, cross-team architecture leadership, strategic influence, and creating conditions for successful collaboration across teams.
This position is based in Redmond, Washington and requires the employee be in the office a minimum of 3 days per week.
Responsibilities
3 Partner with broad teams to design, develop, validate, debug, and optimize infrastructure software across OS, firmware, fleet health, hardware repair, telemetry, and monitoring areas.
3 Define the technical strategy and architecture for New Product Introduction across Microsoft 365 fleet infrastructure, including platform bring-up, validation, deployment readiness, and lifecycle management.
3 Lead hardware and firmware reliability architecture, including health telemetry, diagnostics, failure detection, predictive insights, and remediation automation.
3 Drive failure analysis and root cause investigation for critical live-site incidents involving hardware, firmware, OS, storage, networking, or infrastructure platforms, and drive durable improvements that reduce recurrence.
3 Partner across Microsoft 365, Azure Hardware Systems, silicon vendors, OEM/ODM partners, firmware teams, and operations organizations to improve platform reliability, scalability, and serviceability.
3 Establish engineering standards, observability patterns, and operational mechanisms that improve fleet availability, reduce deployment risk, and strengthen hyperscale infrastructure operations.
3 Evolve M365 substrate infrastructure strategy for AI and agentic workloads across compute, memory, storage, and networking domains.
3 Influence cross-organizational technical direction, communicate complex tradeoffs clearly, and help teams make durable architecture decisions.
3 Mentor early in career engineers.
3 Embody Microsoft's culture and values.
Qualifications
Required Qualifications:
3 Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
3 OR equivalent experience.
Other Requirements:
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:
3 Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
Preferred Qualifications:
3 Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
3 OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
3 OR equivalent experience.
3 Experience leading architecture and technical strategy for large-scale distributed systems, cloud infrastructure, or hyperscale service platforms.
3 Deep experience with hardware platforms, firmware, OS, drivers, datacenter infrastructure software, storage systems, networking, or platform software.
3 Experience with New Product Introduction, platform validation, deployment readiness, lifecycle management, or fleet-scale hardware operations.
3 Experience on hyperscale fleet live site, using telemetry, observability, failure analysis, predictive diagnostics, or automation to improve infrastructure reliability.
3 Experience working with silicon providers, hardware vendors, OEMs, ODMs, firmware teams, or platform engineering organizations.
3 Experience influencing cross-organizational engineering strategy and driving complex technical programs across multiple teams.
3 Communication skills with the ability to explain technical tradeoffs to senior engineering and business leaders.
3 Microsoft Secure environment/US Gov clouds.
#M365Core #FleetAndCapacity #AIInfrastructure #HardwareReliability #CloudInfrastructure
Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.