Senior Software Engineer, Inference

Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of Software Engineering experience, 1-2+ years on LLM inference runtimes
  • Degree in Computer Science or related field
  • Proficient in Go and Python, with C++/CUDA debugging skills
  • Strong understanding of LLM inference internals and best practices
  • Advanced knowledge of Kubernetes architectures
  • Experience with debugging/profiling multi-tier application workloads
  • Familiarity with distributed systems and GPU memory management

Responsibilities

  • Design and implement components of the LLM serving deployment
  • Collaborate with teams to enhance inference efficiency metrics
  • Build distributed execution capabilities for model serving
  • Evaluate and recommend adoption of emerging runtimes and techniques
  • Contribute to orchestration layer for GPU scheduling and cache management
  • Troubleshoot and resolve end-to-end customer issues
  • Lead code reviews and mentor junior engineers

Benefits

  • Comprehensive suite of health and wellbeing benefits
  • Programs for personal and professional development
  • Commitment to inclusivity and support for diverse backgrounds
  • Flexibility to manage work and personal needs
  • Encouragement for career growth and skills application
Full Job Description
Senior Software Engineer, Inference

This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Job Description:

HPE's Private Cloud AI organization is seeking a Senior Software Engineer to build and evolve the model runtime within HPE AI Essentials, the inference platform used by enterprises to operate large language models on infrastructure they own, including air-gapped and sovereign environments. The core engineering challenge in this domain is not model deployment but sustained execution efficiency: achieving low tail latency and high GPU utilization on customer-owned hardware of varying generation and configuration. In this role you will design and implement key components of that runtime - engine integration, batching, KV cache management, and distributed execution - together with the Kubernetes orchestration layer that supports it. The primary work location is as listed, but could be any other HPE site location in the US; however, remote work options will be considered.
Responsibilities
• Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
• Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
• Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
• Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption
• Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
• Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence
• Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team

Knowledge and Skills

Required
• Familiar with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
• Strong understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
• Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them
• Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
• Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight
• Familiar with debugging/profiling multi-tier application workloads such as RAG, Agents
• Excellent analytical, debugging, and problem-solving abilities

Preferred
• Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe
• Disaggregated prefill/decode serving, or KV cache offload and reuse at scale
• RDMA, GPUDirect Storage, InfiniBand, or RoCE
• MIG, fractional GPU allocation, and multi-tenant GPU isolation
• On-premises, air-gapped, or regulated enterprise software delivery

Experience and Education
• Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving
• Degree in Computer Science or related field

Accessibility

HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here.

Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability.

What We Can Offer You:

Health & Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal & Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have - whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

Let's Stay Connected:



#unitedstates

Job:
Engineering
Job Level:
TCP_04

"The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.
- United States of America: Annual Salary USD 144,000 - 273,000 in Colorado // 137,000 - 315,000 in North Carolina & Texas
The listed salary range reflects base salary. Variable incentives may also be offered."

Information about employee benefits offered in the US can be found at https://myhperewards.com/main/new-hire-enrollment.html

The estimated job application period closure is December 30 2027; this timeline is provided for transparency and internal planning purposes.

About Hewlett Packard Enterprise Development LP

Hewlett Packard Enterprise Development LP Careers

Joining Hewlett Packard Enterprise Development LP presents an unparalleled opportunity to advance a career in technology alongside some of the industry's most innovative minds. Hewlett Packard Enterprise Development LP stands as a beacon of innovation, leadership, and professional growth, offering a plethora of job opportunities that cater to a diverse range of skills and experiences.

Explore Career Opportunities

Hewlett Packard Enterprise Development LP is actively hiring, seeking passionate, creative, and solution-driven team players. Explore open positions that align with your skills and interests in areas ranging from engineering to marketing, and sales to IT. Each position at Hewlett Packard Enterprise Development LP not only boosts professional growth but also contributes to the company's culture of innovation and leadership.

Internship Programs

Kickstart your career with Hewlett Packard Enterprise Development LP’s internship programs. These opportunities allow interns to work on real projects, gaining hands-on experience and insights into the company's operations. Internships are a gateway to full-time employment, offering invaluable networking opportunities and a chance to build a professional resume.

Employee Benefits and Culture

Hewlett Packard Enterprise Development LP is committed to fostering a workplace where diversity and inclusion are integral to the company culture. Employees enjoy a range of benefits designed to support their physical, financial, and emotional well-being. From health insurance to retirement plans and flexible working conditions, the company ensures that team members have what they need to succeed.

Professional Development and Growth

The commitment to employee growth is evident through comprehensive training and development programs that encourage continuous learning and career advancement. Leadership development and diversity training are pillars of the company's strategy, ensuring that all team members have the opportunity to lead and innovate.

Networking and Innovation

At Hewett Packard Enterprise Development LP, networking goes hand in hand with innovation. Employees are encouraged to connect with colleagues and leaders through various professional networks and company-sponsored events. This culture of collaboration drives the development of groundbreaking solutions and services, reinforcing the company's position at the forefront of the technology sector.

Join the Hewlett Packard Enterprise Development LP Team

Search for job opportunities that match your skills and interests. Hewlett Packard Enterprise Development LP looks for individuals eager to drive innovation and lead in the digital era. Prepare your resume, hone your interview skills, and get ready to join a team that values growth and leadership.

Stay Connected with Hewlett Packard Enterprise Development LP Careers

Keep up to date with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the professionals who work at Hewlett Packard Enterprise Development LP.

Sign Up for Job Alerts

Personalize your subscription to receive job alerts, latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at Hewlett Packard Enterprise Development LP.
Learn more about Hewlett Packard Enterprise Development LP

Similar Jobs

More Jobs at Hewlett Packard Enterprise Development LP

More Information Technology Jobs

Find similar Senior Software Engineer, Inference jobs: