You will:
- Work with AI agents to create high performance low-level kernels in a Domain-Specific language for a custom ML accelerator
- Ensure functional correctness and numerical accuracy of kernels
- Conduct performance analysis of the kernels to identify performance bottlenecks and areas that need improvement
- Develop a custom harness to recursively improve kernel generation and remove the need for human involvement
You have:
- Currently enrolled in an MS or PhD program in Computer Science, Computer Engineering, Computational Science, or a related quantitative field
- Demonstrated track record of using AI agents and harnesses to improve coding performance
- Working knowledge of key kernels for transformer inference including performance analysis and optimization
- Familiarity with one or more low-level ML accelerator DSLs (CUDA, Triton, Gluon, etc.)
We prefer:
- Solid understanding of modern inference accelerator architectures
- Fluency in Python and/or C++
- Working knowledge of different ML numerical formats and their relative accuracy vs performance tradeoffs
Note: This will be a hybrid onsite internship position. We will accept resumes on a rolling basis until the role is filled. To be in consideration for multiple roles, you will need to apply to each one individually - please apply to the top 3 roles you are interested in.
The expected hourly rate for this full-time position is listed below. Interns are also eligible to participate in the Company's generous benefits programs, subject to eligibility requirements.
Hourly Masters Pay
$70-$70 USD
The expected hourly rate for this full-time position is listed below. Interns are also eligible to participate in the Company's generous benefits programs, subject to eligibility requirements.
Hourly PhD Pay
$85-$85 USD