F5 Networks

AI Inference Engineer

F5 Networks$176K — $265K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years experience in high-performance AI inference roles.
  • Proficient in Python, C++, Rust, or Golang for AI applications.
  • Hands-on experience with inference tools like vLLM, TensorRT, and Llama.cpp.
  • Familiarity with Docker, Kubernetes, and major cloud platforms (AWS, GCP, Azure).
  • Solid understanding of GPU and AI hardware programming and optimization techniques.

Responsibilities

  • Build and maintain high-performance inference engines with vLLM, TGI, and NVIDIA Triton.
  • Deliver low-latency AI serving solutions across multiple applications.
  • Optimize models for specialized hardware like NVIDIA GPUs and TPUs.
  • Collaborate with hardware teams for maximum performance across computing environments.
  • Design auto-scaling architectures for real-time and batch inference in Kubernetes.
  • Monitor performance metrics against SLAs to maintain system reliability.
  • Execute performance tests to identify and resolve bottlenecks.

Benefits

  • Generous health and wellness plans tailored to employee needs.
  • Flexible work environment that promotes work-life balance.
  • Professional development opportunities in advanced AI techniques.
  • Collaborative workplace culture that emphasizes team innovation.
Full Job Description
Job Description

The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments-from GPU-rich data centers to resource-constrained edge devices-with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy.

This role is pivotal in advancing F5's AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance.

Key Responsibilities

High-Performance AI Serving
  • Build and maintain robust inference engines using tools like vLLM, TGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale.
  • Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications.

Hardware Acceleration and Optimization
  • Profile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs.
  • Collaborate with hardware teams to maximize utilization and performance across various computational environments.

Inference Orchestration and Scalability
  • Design and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration.
  • Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability.

Performance Monitoring and Observability
  • Establish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs).
  • Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale.


Technical Requirements

Required Skills:
  • Programming Languages: Proficiency in programming languages such as Python, C++, Rust, or Golang specifically for high-performance AI workflows.
  • Inference Tools: Proven hands-on experience with tools like vLLM, TensorRT, Llama.cpp, and Ollama for inference development and optimization.
  • Infrastructure Expertise: Strong familiarity with infrastructure technologies, including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure.
  • Hardware Optimization Expertise: Comprehensive understanding of GPU and AI hardware, including techniques for profiling and optimizing performance for accelerators like NVIDIA GPUs and TPUs.


Preferred Experience:
  • Prior experience deploying Large Language Models (LLMs) with advanced techniques like Speculative Decoding or PagedAttention.
  • Contributions to open-source inference libraries or hardware-level kernel development (e.g., CUDA, Triton kernels).
  • Background in MLOps or SRE roles focused on high-performance AI endpoints and reliability during demand surges.
  • Proficiency in designing scalable solutions for high-throughput inference environments optimized for traffic bursts.


Success Metrics (KPIs):
  • Latency Reduction: Continuously improve inference latency metrics, ensuring minimal Time to First Token (TTFT) and maximum tokens per second.
  • Cost Efficiency: Achieve lower "Cost per 1K Tokens" through better resource utilization and hardware optimization.
  • Scalability: Maintain system stability and reliability during traffic spikes, ensuring performance consistency across environments.
  • Throughput Maximization: Deploy models optimized for peak hardware usage and maximized process throughput.


The base pay range per annum for this position is: $176,600 - $265,000

F5 maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, geographic locations, and market conditions, as well as to reflect F5's differing products, industries, and lines of business. The pay range referenced is as of the time of the job posting and is subject to change. You may also be offered incentive compensation, bonus, restricted stock units, and benefits. More details about F5's benefits can be found at the following link: https://www.f5.com/company/careers/benefits. F5 reserves the right to change or terminate any benefit plan without notice.

#LI-ZB1

The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.

Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com or @myworkday.com)

About F5 Networks

F5 Networks, Inc. is a global company that specializes in application services and application delivery networking (ADN). F5 technologies focus on the delivery, security, performance, and availability of web applications, as well as the availability of servers, cloud resources, data storage devices, and other networking components. F5 Networks is a leading provider of application delivery networking (ADN) technology that optimizes the delivery of network-based applications and the security, performance, and availability of servers, data storage devices, and other networking components. F5 Networks is a leading provider of application delivery networking (ADN) technology that optimizes the delivery of network-based applications and the security, performance, and availability of servers, data storage devices, and other networking components. F5 Networks is a leading provider of application delivery networking (ADN) technology that optimizes the delivery of network-based applications and the security, performance, and availability of servers, data storage devices, and other networking components.
Learn more about F5 Networks
Size
6,461 employees
Market Cap
$8.4 billion
Industry
Net Income
$296.5 million
Founded
1996
5 Year Trend
+5.2%
Revenue
$2.4 billion
NASDAQ

Similar Jobs

More Jobs at F5 Networks

More Enterprise Technology Jobs

Find similar AI Inference Engineer jobs: