Minimum qualifications:- Bachelor's degree or equivalent practical experience.
- 8 years of experience in software development.
- 5 years of experience in programming, debugging, computer architecture, and C .
- 3 years of experience in a technical leadership role.
- Experience in a people management, supervision/team leadership role.
Preferred qualifications:- Master's degree or PhD in Engineering, Computer Science, or a related technical field.
- Experience with Python, standard ML, high-performance computing, and embedded systems.
- Background in machine learning, accelerator architectures (driver/run times), or related fields.
About the jobWith technical and leadership expertise, you manage engineers across multiple teams and locations, a large product budget and oversee the deployment of large-scale projects across multiple sites internationally.
In this role, you will have unique and exciting opportunity to work at the "heart of machine learning" and learn about basic to advanced ML paradigms for training/inference while delivering impact across different compute infrastructure, Cloud, and Open Source environments. You will focus on hardware/software interactions with accelerators (current and chips under design) and integration with the rest of the profiling system.
Responsibilities- Be focused around driving continuous improvements to the machine learning software/hardware stacks through providing insightful performance debugging. Provide insights by summarizing different views of captured profile data such as trace timelines, memory usage, HLO profiles, ML graph summaries.
- Learn and build an intuitive understanding of existing data collection, analysis, and visualization workflows.
- Support new and exciting ML paradigms (such as horizontal scaling for upcoming TPU chips) by making contributions across the end to end stack and analysis tools.
- Partner with product area leads to understand model optimization use cases, drive cross functional efforts to deliver on chip profiling requirements, and propose new hardware features.
- Collaborate across Hardware, Driver, Runtime, and Performance Analysis teams and many other stakeholders.