Senior/Staff AI Engineer

Data Direct Networks

$150K — $180K *
US-AnywhereRemote in California, US
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience building or optimizing production AI systems.
  • Strong understanding of inference performance across compute, memory, and storage architecture.
  • Hands-on experience at the systems layer, specifically with GPU and CPU resource management.
  • Demonstrated ownership in model serving, retrieval, caching, storage, or distributed performance.
  • Ability to bridge architectural decisions and hands-on implementation in efficiency-driven environments.
  • Background in technically demanding fields like AI infrastructure or high-performance systems is essential.
  • PhD preferred but real-world experience is more valuable.

Responsibilities

  • Build and optimize LLM serving and inference systems for production.
  • Enhance performance on GPU and CPU pathways.
  • Address KV cache, memory, storage, and throughput bottlenecks.
  • Design and scale systems for retrieval-heavy AI workloads.
  • Contribute to infrastructure that improves AI performance through storage architecture efficiencies.
  • Resolve engineering challenges at the intersection of AI, high-performance systems, and distributed infrastructure.

Benefits

  • Work on the infrastructure that ensures AI systems are fast and scalable.
  • Engage with deep systems problems at the mechanical level of AI.
  • Opportunity to impact how AI performance translates commercially.
  • Collaborate with a team prioritizing performance, scale, and architecture in AI.
  • Be part of cutting-edge advancements in AI infrastructure.
Full Job Description
What you'll do
  • Build and optimize LLM serving and inference systems for production environments
  • Improve performance across GPU and CPU pathways
  • Work on KV cache, memory, storage, and throughput bottlenecks
  • Design and scale systems that support RAG and retrieval-heavy AI workloads
  • Contribute to infrastructure where storage architecture and systems efficiency materially affect AI performance
  • Solve engineering problems at the intersection of AI, high-performance systems, and distributed infrastructure


What we're looking for
  • An engineer who has spent meaningful time building or optimizing production AI systems, not just experimenting with models
  • Someone who understands how inference performance is shaped by the interaction between compute, memory, storage, and serving architecture
  • Deep hands-on experience working close to the systems layer - for example, improving how workloads run across GPU and CPU resources, reducing bottlenecks, or tuning infrastructure for better throughput and latency
  • Evidence of real ownership in areas like model serving, retrieval, caching, storage, or distributed performance, rather than purely application-layer AI work
  • The ability to move comfortably between architecture decisions and hands-on implementation, especially in environments where efficiency and scale matter
  • A background that suggests you can operate in technically demanding environments, whether that comes from AI infrastructure, high-performance systems, storage platforms, or adjacent distributed systems work
  • PhD preferred, but far less important than having built serious systems in the real world


Why this role is compelling
  • This is not a "prompt engineering" job.
  • This is not an "AI wrapper" job.
  • This is not a generic backend role with AI sprinkled on top.
  • This is a chance to work on the infrastructure that determines whether modern AI systems are fast, scalable, efficient, and commercially viable.
  • If you want to work on the real mechanics of AI performance - serving, retrieval, compute efficiency, memory behavior, storage architecture, and inference at scale - this is where that work happens.


Who will love this role
  • Engineers who enjoy deep systems problems
  • Builders who care about performance, scale, and architecture
  • People who want to work where AI meets infrastructure
  • Candidates who would rather solve hard technical bottlenecks than ship surface-level AI features


Who should not apply

This role is not for:
  • Purely academic researchers without meaningful production ownership
  • Generic software engineers without clear AI systems or inference depth
  • Candidates focused mainly on prompt engineering or lightweight application integrations
  • MLOps generalists who have not worked deeply on serving, storage, or performance-critical AI systems

Similar Jobs

More Information Technology Jobs

Find similar Senior/Staff AI Engineer jobs: