Contribute to the development of cutting-edge deep learning kernels for advanced architectures. Collaborate with cross-functional teams to optimize performance and ensure timely delivery of solutions in a fast-paced environment.
Optimize and develop high-performance Tensor Core-based deep learning kernels, collaborating across teams to enhance GPU functionality and efficiency within architecture initiatives for cutting-edge technologies.
Grow your career with a cutting-edge team specializing in GPU architecture and compiler technologies, fostering high-performance kernel design and collaboration for innovative solutions within the MLIR ecosystem.
Accelerate advancements in GPU design by developing cutting-edge Tensor Core abstractions and high-performance kernels in MLIR, Python, and C++. Collaborate with experts to shape the future of computing hardware and software integration.
Join a team that's shaping the future of GPU software through the design and implementation of high-performance kernels and compiler technologies, driving innovation in machine learning and parallel computing.