The Product: AWS Machine Learning accelerators are at the forefront of AWS innovation. Trainium delivers best-in-class ML training performance with the most teraflops (TFLOPS) of compute power for ML in the cloud. This is all enabled by the AWS Neuron Software Development Kit (SDK), which includes an ML compiler, the Neuron Kernel Interface (NKI) compiler, and a runtime that natively integrates into popular ML frameworks such as PyTorch and JAX.
Neuron Kernel Interface (NKI) is a bare-metal language and compiler for directly programming NeuronDevices available on AWS Trainium instances. You can use NKI to develop, optimize, and run new operators directly on NeuronCores while making full use of available compute and memory resources.
Explore NKI:
- https://awsdocs-neuron.readthedocs-hosted.com/en/latest/nki/index.html
AWS Neuron is used at scale by customers such as Epic Games, Snap, Airbnb, Autodesk, Amazon Alexa, and Amazon Rekognition, along with many others across a range of segments.
You: We are seeking a talented Software Development Manager with strong leadership and mentoring skills to join our NKI development team. As an SDM III, you will lead a team of experienced compiler engineers developing compiler optimization algorithms and deploying, at scale, a new compiler targeting AWS custom hardware. You will need to be technically capable, credible, and curious in your own right as a trusted AWS Neuron manager, innovating on behalf of our customers. You will draw on knowledge of resource management, scheduling, code generation, optimization, and instruction architectures across CPU, NPU, GPU, and novel forms of compute. You will leverage your technical communication skills to partner with AWS ML services teams and pre-silicon design, and to bring new products and features to market. As deep learning models become more versatile, using compiler technologies to achieve both high performance and high productivity becomes essential. Join the team to build the software that boosts the entire deep learning community.
Explore the Product:
- https://aws.amazon.com/machine-learning/neuron/
- https://awsdocs-neuron.readthedocs-hosted.com/en/latest/general/nki/index.html
In order to be considered for this role, candidates must be currently located in or willing to relocate to Cupertino, CA.
Key job responsibilities
- Lead, grow, and mentor a team of compiler engineers, including hiring, career development, and team health
- Own execution and delivery of NKI compiler features across release cycles, balancing scope, quality, and timelines
- Set technical direction and roadmap for your area of the NKI compiler in partnership with senior engineers
- Partner across teams, including frameworks, kernel development, runtime, hardware/pre-silicon design, and product management, to bring new features to market
- Drive resolution of complex, high-priority technical issues, staying close enough to the code and hardware to guide the team
- Represent your team's work and priorities to senior leadership and cross-organizational stakeholders
- Anchor decisions in customer impact, ensuring the team builds what unblocks real ML workloads on Trainium
A day in the life
No two days look the same on the NKI team, but most blend people leadership, technical depth, and cross-team collaboration.
You might start the morning in a 1:1 with an engineer, working through the design of a new compiler optimization or unblocking a tricky scheduling problem. Mid-morning, you join a release sync to check in on branch-cut readiness and make a call on what makes the current train versus the next. After lunch, you catch up with a partner team, kernel developers, runtime, or hardware design, to align on an upcoming feature and its dependencies.
In the afternoon, you spend focused time on the roadmap: shaping where NKI is heading over the next few quarters, and translating customer needs into concrete technical investments. You review a design doc, leave feedback that sharpens the team's thinking, and dig into a profiling result yourself to understand where a real workload is leaving performance on the table. You close the day by clearing a path for your team and resolving an escalation, connecting two people who should be talking, or writing up a decision so the team can move fast tomorrow.
Throughout, you keep one question at the center: what does this unblock for our customers? You are technically credible enough to earn your team's trust, and you spend your energy on the highest-leverage problems: growing your people, delivering the compiler, and inventing on behalf of the customers who run their most demanding ML workloads on Trainium.
BASIC QUALIFICATIONS
- 2+ years of engineering team management experience
- 6+ years of working directly within engineering teams experience
- 4+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
- Experience partnering with product or program management teams
- Understanding of compilers (resource management, instruction scheduling, code generation, and compute graph optimization)
- Strong software design fundamentals and excellent system-level coding skills
PREFERRED QUALIFICATIONS
- M.S. or Ph.D. in Computer Science or related technical field
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, CA, Cupertino - 212,700.00 - 287,700.00 USD annually