Innovate high-performance kernel design for Tensor Core and GPU features. Engage with the MLIR backend compiler stack and collaborate with architecture teams to enhance GPU performance. Drive advancements in CUDA C++ and CUTLASS Python DSL.
Champion the future of semiconductor testing by defining and leading innovative technical strategies. Collaborate with cross-functional teams to enhance testing capabilities and ensure scalability for complex device portfolios.
Unlock your potential by benchmarking and optimizing deep learning performance on GPUs. Collaborate across teams to enhance CUTLASS and lead the way in software performance improvements and innovative tooling.
Optimize high-performance computations by benchmarking deep learning models and enhancing GPU kernel efficiency. Collaborate with teams to close performance gaps and develop automation tools for continuous optimization within the CUTLASS ecosystem.
Engage with cross-functional teams to benchmark, analyze, and optimize deep learning model performance using CUTLASS, driving innovation in GPU computing and AI workflows.