Job DescriptionAt LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. This role may be remote or hybrid. At LinkedIn, hybrid roles are performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team. Remote roles are performed from the designated home work location upon time of hire, and any changes to this home work location requires a review of remote status and approval.
We're hiring a Principal Staff Software Engineer to lead LinkedIn's GPU-Based Retrieval Platform, a foundational AI infrastructure stack that powers candidate generation and retrieval across Feed, Ads, Search, Talent, and other critical product experiences. The platform sits on the hot path of billions of member interactions each day, making this one of the highest-leverage technical leadership roles within LinkedIn's AI Infrastructure organization.
In this role, you will own the platform's technical direction and architecture end to end, spanning large-scale indexing and retrieval, low-latency distributed serving, GPU scheduling, memory efficiency, batching, and kernel-level optimization. You will drive improvements in throughput, tail latency, retrieval quality, reliability, and cost, directly influencing member engagement, business outcomes, and AI engineer productivity across the company.
The GPU-Based Retrieval Platform team works at the intersection of GPU systems, distributed serving, information retrieval, machine learning, and product engineering. You will partner closely with teams across Feed, Ads, Search, Talent, Modeling, and Infrastructure, while setting the technical direction for a multi-team platform, influencing cross-company architecture, and mentoring senior engineers.
Responsibilities:- Set the long-term technical strategy and architecture for LinkedIn's GPU-Based Retrieval Platform.
- Lead the design and evolution of large-scale indexing, candidate generation, vector search, and retrieval-serving systems.
- Optimize GPU performance across CUDA, Triton, memory hierarchies, batching, scheduling, and multi-GPU communication.
- Improve throughput, QPS per GPU, tail latency, recall quality, reliability, and infrastructure cost.
- Build scalable, observable, and highly available multi-tenant serving systems for high-QPS production workloads.
- Make critical architectural trade-offs across latency, quality, capacity, model complexity, and cost.
- Partner with Feed, Ads, Search, Talent, Modeling, and Infrastructure teams to shape the platform roadmap.
- Evaluate emerging GPU technologies, retrieval architectures, and serving frameworks for adoption at LinkedIn.
- Lead complex cross-organizational initiatives from architecture and design through production rollout and adoption.
- Mentor senior engineers, raise the technical bar, and influence AI infrastructure strategy across LinkedIn.
QualificationsBasic Qualifications:- BS in Computer Science or equivalent.
- 10+ years of industry experience in software design, development, and algorithm related solutions.
- 5+ years in experience as an architect, or technical leadership position.
- Experience in developing and scaling large scale databases or analytics systems
- Experience with coding in Java, C++, or Rust; knowledge of query execution, indexing, and concurrency.
- Hands on experience developing distributed systems, large-scale systems, databases and/or Backend APIs
Preferred Qualifications:- Master's or PhD in Computer Science or a related technical discipline, with experience operating at Principal Staff or equivalent scope.
- 15+ years of software engineering experience, including 7+ years in senior technical leadership roles shaping architecture across multiple organizations.
- 5+ years of hands-on experience with CUDA, Triton, GPU kernel optimization, and hardware-aware performance tuning.
- 3+ years of experience with NCCL, distributed inference, multi-GPU communication, or optimizing workloads on modern accelerators such as NVIDIA H100 or H200 GPUs.
- 3+ years of experience with inference optimization techniques such as quantization, mixed precision, batching, memory management, and throughput or latency tuning.
- 5+ years of experience building large-scale retrieval systems, including ANN algorithms, hybrid retrieval, learned indexes, or billion-scale vector search.
- 5+ years of experience building multi-tenant AI serving or retrieval platforms and balancing recall, latency, throughput, reliability, and cost across multiple products, models, or workloads.
- Experience in one or more of the following domains: search, recommendations, feed, advertising, candidate generation, LLM serving, MLOps, or large-scale AI infrastructure.
- Hands-on experience with one or more of the following: CUDA, Triton, GPU scheduling, memory optimization, multi-GPU workloads, embeddings, vector search, ANN, or candidate generation.
Suggested Skills:- AI / ML Infrastructure
- Technical Strategy
- Distributed Systems
- Stakeholder Management
LinkedIn is committed to fair and equitable compensation practices.
The pay range for this role is $207,000 to $340,000. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor.
The total compensation package for this position may also include annual performance bonus, stock, benefits and/or other applicable incentive compensation plans. For more information, visit https://careers.linkedin.com/benefits.
Additional Information