OverviewThe AI Infrastructure team is responsible for building and operating large-scale, highly reliable, and efficient GPU infrastructure that powers Microsoft's AI ecosystem. We host the training and inference platforms behind many of Microsoft's flagship AI offerings, including Microsoft 365 Copilot, GitHub Copilot, Microsoft Copilot, and Azure AI Foundry's inference and fine-tuning services for both OpenAI and open-source models. Our infrastructure enables AI innovation at hyperscale and supports some of the most demanding workloads across the company.
As a Software Engineer on the AI infrastructure team, you will work on cutting edge infrastructure and tools to support large scale model deployments, pre-training, post-training and fine-tuning on latest generation of NVIDIA and AMD GPUs in Azure and Microsoft partner clouds on some of the world's largest AI Supercomputers.
ResponsibilitiesAs an engineer on the AI infrastructure team, your responsibilities include:
• Design, develop, and maintain AI infrastructure services in Go, Rust, Python, C++, and C#, deployed on large-scale Kubernetes clusters to support inference, pre-training, and post-training workloads for state-of-the-art AI models.
• Collaborate with engineers, researchers, and external partners to troubleshoot issues, improve reliability, and optimize the performance of large-scale AI training and inference systems.
• Build and enhance distributed systems that deliver high reliability, low latency, operational efficiency, and strong security across Azure and partner cloud environments.
• Develop automation and tooling to improve GPU capacity utilization, streamline fleet operations, and enable efficient scaling of AI infrastructure.
• Provide operational support, technical leadership, and vision while contributing to the deployment, monitoring, and continuous improvement of engineering systems and practices.
QualificationsRequired Qualifications:- Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
Preferred Qualifications: • 2+ years designing, developing, and shipping high quality software.
• 2+ years of experience with distributed systems and cloud-based infrastructure.
• 1+ year of experience with DevOps practices (CI/CD, automated testing, deployment, etc.).
• 2+ years of software development experience in C#, C++, Python, or similar languages.
• 2+ years of experience with containerization tools (e.g., Docker, Kubernetes).
• Knowledge and hands on experience with production ML systems, large-scale training infrastructure, NCCL, CUDA libraries and tools
#AIINFRA
Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.