At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.
LinkedIn's AI Infrastructure organization is responsible for building the foundational platforms that power AI across LinkedIn. The LLM Serving team builds the critical infrastructure that enables efficient, reliable, and large-scale deployment of large language models and other advanced AI models in production.
This team sits at the center of LinkedIn's AI platform, owning the layer between model training and production serving. The work focuses on making large-scale models run faster, cheaper, and more efficiently on GPUs at LinkedIn scale. The team builds and extends high-performance serving infrastructure and contributes to leading open-source technologies such as SGLang, vLLM, and related model serving frameworks.
We are looking for a Principal Staff Software Engineer with deep expertise at the intersection of systems, machine learning, GPU infrastructure, and large-scale inference. This is a highly technical, company-level leadership role for an engineer who can set long-term technical direction while remaining deeply hands-on across the serving stack.
You will help define the architecture and evolution of LinkedIn's next-generation LLM serving platform, driving improvements in performance, efficiency, reliability, scalability, and cost across AI workloads. The role requires the ability to operate across model architecture, runtimes, compilers, kernels, distributed systems, and hardware while influencing technical strategy across multiple teams and organizations.
Responsibilities- Set the long-term technical strategy and architecture for LinkedIn's large-scale LLM serving and inference infrastructure.
- Lead the design, development, and evolution of high-performance online and offline inference platforms for LLMs and other advanced AI models.
- Drive major improvements in inference latency, throughput, GPU utilization, reliability, scalability, and infrastructure cost.
- Architect serving systems that operate efficiently across large GPU fleets and support a diverse set of models, products, and production workloads.
- Optimize model execution across the full stack, including model architecture, serving runtime, compiler, kernel, memory, networking, and hardware layers.
- Drive adoption of model optimization techniques such as quantization, pruning, compression, batching, caching, and memory optimization.
- Improve GPU efficiency through low-level systems work, including CUDA and Triton optimization, kernel-level improvements, runtime tuning, scheduling, and hardware-aware performance engineering.
- Make critical architectural trade-offs across latency, throughput, model quality, capacity, reliability, developer experience, and cost.
- Partner with ML, infrastructure, product, and research teams to identify systemic serving bottlenecks and shape the long-term AI infrastructure roadmap.
- Evaluate and drive adoption of emerging inference technologies, serving architectures, accelerators, and open-source frameworks.
- Contribute to and extend open-source LLM serving technologies such as SGLang, vLLM, Triton, TensorRT, Ray, or similar frameworks.
- Lead complex cross-organizational initiatives from architecture and design through production rollout, adoption, and operational maturity.
- Mentor senior technical leaders, raise the engineering bar, and influence AI infrastructure strategy across LinkedIn.
QualificationsBasic Qualifications- BA/BS degree in Computer Science or a related technical field, or equivalent practical experience.
- 10+ years of industry experience in software engineering, distributed systems, infrastructure, machine learning systems, or related technical areas.
- 5+ years of experience in an architect, technical lead, or senior technical leadership capacity, driving architecture and technical direction across complex systems or multiple engineering teams.
- Experience designing, building, and scaling large-scale production ML systems, model serving platforms, AI infrastructure, or distributed systems.
- Experience building or optimizing GPU-based inference systems, including CUDA, Triton, kernel optimization, runtime optimization, or hardware-aware performance tuning.
- Hands-on programming experience in one or more languages such as C++, Python, Go, Java, or Rust, with experience leading technically complex initiatives across team or organizational boundaries.
Preferred Qualifications- Master's or PhD in Computer Science, Machine Learning, Electrical Engineering, or a related technical field, with experience operating at Principal Staff or equivalent scope.
- 15+ years of software engineering experience, including 7+ years in senior technical leadership roles shaping architecture across multiple teams or organizations.
- 5+ years of hands-on experience building or optimizing large-scale LLM serving, AI inference, GPU infrastructure, or production model deployment platforms.
- Deep experience with CUDA, Triton, GPU kernel optimization, distributed inference, multi-GPU systems, or hardware-aware performance tuning.
- Experience with inference optimization techniques such as quantization, mixed precision, batching, caching, memory optimization, and latency or throughput tuning.
- Experience building large-scale, multi-tenant AI serving platforms and balancing performance, reliability, scalability, GPU utilization, and infrastructure cost.
- Familiarity with or contributions to serving and runtime frameworks such as vLLM, SGLang, Triton, TensorRT, Ray, XLA, TVM, or similar technologies.
- Demonstrated ability to set long-term technical strategy, lead complex cross-organizational initiatives, and influence architecture across ML, infrastructure, research, and product teams.
Suggested Skills:- AI / ML Infrastructure
- Technical Strategy
- Distributed Systems
- Stakeholder Management
LinkedIn is committed to fair and equitable compensation practices.
The pay range for this role is $207,000 to $340,000. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor.
The total compensation package for this position may also include annual performance bonus, stock, benefits and/or other applicable incentive compensation plans. For more information, visit https://careers.linkedin.com/benefits.