Your role and responsibilities
Join a team conducting cutting-edge research in AI Native Distributed Systems. In this role, you will explore and advance one or more of many aspects of AI-native distributed computing we are currently working on, including: design, development, and optimization of software platforms for building, deploying, and managing large language models (LLMs) and agents; principles, methodologies, and frameworks for AI-native software development, evolution, and optimization; high-performance distributed inference for LLMs; frameworks, tools, and infrastructure that support the end-to-end AI lifecycle;
networking for AI infrastructure, developing/prototyping networking solutions and techniques for future AI systems, including networking for distributed inference and RDMA over Ethernet networking, as well as evaluation of the solutions from performance, scalability, and resiliency perspectives; networking innovations that enable high-performance communication for AI workloads and emerging distributed inference patterns.
The ideal candidate has a strong foundation in AI, large language models, and agentic workloads, combined with expertise in systems, networking, or software engineering. As an intern, you will contribute across the full industrial research lifecycle: formulating novel ideas, designing and building prototype systems, evaluating performance at scale, demonstrating real-world impact, and publishing results in leading scientific venues.
Required education
Bachelor's Degree
Preferred education
Master's Degree
Required technical and professional expertise
- Currently pursuing a degree in Computer Science, Engineering, or related field (MS/PhD), with a focus on artificial intelligence, systems, networking, or software development.
- Background in artificial intelligence and machine learning, including areas such as deep learning, LLMs, distributed LLM inference, or agentic patterns
- Strong background in software development, including proficiency in languages such as Python, Go, C++, or Rust
- Experience with containerization using Docker, Kubernetes, or other container orchestration tools.
Preferred technical and professional experience
- Understanding of LLM technology, distributed inference, llm-d/vLLM
- Designing and implementing distributed systems
- Experience in cloud/data Center networking, Software-Defined Networking (SDN), network virtualization, or Linux networking
- Research and development experience in AI platforms, including the design, implementation, and optimization of AI frameworks and tools
OTHER RELEVANT JOB DETAILSSupplemental 1 employees may be eligible for up to 8 paid holidays, minimum of 56 hours paid sick time and the IBM Employee Stock Purchase Plan. IBM offers paid family medical leave and disability benefits to eligible employees where required by applicable law.
This position was posted on the date cited in the key job details section and is anticipated to remain posted for 15 days from this date or less if not needed to fill the role.
We consider qualified applicants with criminal histories, consistent with applicable law.
IBM will not be providing visa sponsorship for this position now or in the future. Therefore, in order to be considered for this position, you must have the ability to work without a need for current or future visa sponsorship.
The compensation range and benefits for this position are based on a full-time schedule for a full calendar year. The salary will vary depending on your job-related skills, experience and location. Pay increment and frequency of pay will be in accordance with employment classification and applicable laws. For part time roles, your compensation and benefits will be adjusted to reflect your hours. Benefits may be pro-rated for those who start working during the calendar year.