Sr. Staff Software Engineer, AI Infrastructure

LinkedIn

$198K — $326K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • BS/BA in Computer Science or related field, or equivalent experience
  • 5+ years in software design, development, and algorithms
  • Extensive programming experience in languages such as Python, C++, or Java
  • 2+ years in an architectural or technical leadership role
  • Experience in creating deep learning systems
  • Hands-on experience with distributed or large-scale systems

Responsibilities

  • Own the technical strategy for complex AI challenges
  • Design and optimize large-scale distributed training and serving systems
  • Enhance system observability and developer productivity
  • Mentor engineers and shape the technical culture
  • Collaborate with the open-source community on cutting-edge projects
  • Serve as tech lead for multiple key AI Infrastructure initiatives

Benefits

  • Hybrid work model with flexible scheduling
  • Opportunities for mentorship and professional growth
  • Access to cutting-edge AI technology and projects
  • Involvement with the open-source community
  • Working in a collaborative and innovative team environment
Full Job Description
Job Description

This role will be based in Mountain View, CA or Bellevue, WA.

At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.

Join us to push the boundaries of scaling large models together. The team is responsible for scaling LinkedIn's AI model training, feature engineering and serving with hundreds of billions of parameters models and large scale feature engineering infra for all AI use cases from recommendation models, large language models, to computer vision models. We optimize performance across algorithms, AI frameworks, data infra, compute software, and hardware to harness the power of our GPU fleet with thousands of latest GPU cards. The team also works closely with the open source community and has many open source committers (TensorFlow, Horovod, Ray, vLLM, Hugginface, DeepSpeed etc.) in the team. Additionally, this team focussed on technologies like LLMs, GNNs, Incremental Learning, Online Learning and Serving performance optimizations across billions of user queries.

Model Training Infrastructure: As an engineer on the AI Training Infra team, you will play a crucial role in building the next-gen training infrastructure to power AI use cases. You will design and implement high performance data I/O, work with open source teams to identify and resolve issues in popular libraries like Huggingface, Horovod and PyTorch, enable distributed training over 100s of billions of parameter models, debug and optimize deep learning training, and provide advanced support for internal AI teams in areas like model parallelism, tensor parallelism, Zero++ etc. Finally, you will assist in and guide the development of containerized pipeline orchestration infrastructure, including developing and distributing stable base container images, providing advanced profiling and observability, and updating internally maintained versions of deep learning frameworks and their companion libraries like Tensorflow, PyTorch, DeepSpeed, GNNs, Flash Attention. PyTorch Lightning and more and more.

Model Serving Infrastructure: this team builds low latency high performance applications serving very large & complex models across LLM and Personalization models. As an engineer, you will build compute efficient infra on top of native cloud, enable GPU based inference for a large variety of use cases, cuda level optimizations for high performance, enable on-device and online training. Challenges include scale (10s of thousands of QPS, multiple terabytes of data, billions of model parameters), agility (experiment with hundreds of new ML models per quarter using thousands of features), and enabling GPU inference at scale.

As a Sr. Staff Software Engineer, you will have first-hand opportunities to advance one of the most scalable AI platforms in the world. At the same time, you will work together with our talented teams of researchers and engineers to build your career and your personal brand in the AI industry.

Responsibilities:
  • Owning the technical strategy for broad or complex requirements with insightful and forward-looking approaches that go beyond the direct team and solve large open-ended problems.
  • Designing, implementing, and optimizing the performance of large-scale distributed serving or training for personalized recommendation as well as large language models.
  • Improving the observability and understandability of various systems with a focus on improving developer productivity and system sustenance.
  • Mentoring other engineers, defining our challenging technical culture, and helping to build a fast-growing team.
  • Working closely with the open-source community to participate and influence cutting edge open-source projects (e.g., vLLMs, PyTorch, GNNs, DeepSpeed, Huggingface, etc.).
  • Functioning as the tech-lead for several concurrent key initiatives AI Infrastructure and defining the future of AI Platforms.


Qualifications

Basic Qualifications:
  • BS/BA in Computer Science or related technical field or equivalent technical experience
  • 5+ years of industry experience in software design, development, and algorithm related solutions
  • 5+ years of experience programming in object-oriented languages such as Python, C++, Java, Go, Rust, Scala
  • 2+ years of experience as an architect, or technical leadership position
  • 5+ years of experience in the industry with leading / building deep learning systems
  • Hands-on experience developing distributed systems or other large-scale systems

Preferred Qualifications:
  • MS or PhD in Computer Science or related technical discipline.
  • 10+ years of experience in software design, development, and algorithm related solutions with at least 5 years of experience in a technical leadership position
  • 10+ years of experience in an object-oriented programming language such as Python, C++, Java, Go, Rust, Scala
  • 5+ years of experience with large-scale distributed systems and client-server architectures
  • Experience building ML applications, LLM serving, GPU serving.
  • Co-author or maintainer of any open-source projects
  • Expertise in machine learning infrastructure, including technologies like MLFlow, Kubeflow and large scale distributed systems
  • Expertise in deep learning frameworks and tensor libraries like PyTorch, Tensorflow, JAX/FLAX

Suggested Skills:
  • ML Algorithm Development
  • Machine Learning and Deep Learning
  • Information retrieval / recommendation systems
  • Technical leadership

LinkedIn is committed to fair and equitable compensation practices.

The pay range for this role is $198,000 to $326,000. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor.

The total compensation package for this position may also include annual performance bonus, stock, benefits and/or other applicable incentive compensation plans. For more information, visit https://careers.linkedin.com/benefits.

Additional Information

Similar Jobs

More Jobs at LinkedIn

More Information Technology Jobs

Find similar Sr. Staff Software Engineer, AI Infrastructure jobs: