Senior Software Engineer I - AI Inference Data Plane

DigitalOcean

$139K — $174K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in software engineering with a strong focus on AI/ML applications.
  • Expertise in GoLang or Python, including familiarity with gRPC.
  • Hands-on experience with large-scale language models and inference engines like vLLM and TensorRT.
  • Familiarity with distributed inference serving frameworks such as llm-d and NVIDIA Dynamo.
  • Demonstrated contributions to open-source projects related to AI infrastructure.
  • Proficient knowledge of architecture optimization techniques for large language models.
  • Strong understanding of cloud operations and experience in high-scale environments.

Responsibilities

  • Drive the design and development of critical data plane components for AI inference.
  • Architect and improve high-scale, multi-tenant AI inference solutions.
  • Optimize performance of distributed inference hosting with advanced techniques.
  • Collaborate with cross-functional teams to align technical goals with customer needs.
  • Build and maintain Kubernetes-native distributed inference frameworks for model deployment.
  • Address and solve unique distributed-systems challenges related to LLM serving.
  • Contribute to open-source communities by sharing advancements and improvements in AI infrastructure.

Benefits

  • Remote work flexibility.
  • Opportunity to contribute to open-source projects.
  • Culture focused on mentorship and technical excellence.
  • Access to cutting-edge AI technologies and frameworks.
  • Participation in a growing team dedicated to innovation and performance.
Full Job Description
DigitalOcean is expanding its AI Infrastructure layer to support the next generation of AI-driven applications. We are seeking a Senior Engineer 2 to join our AI Inference Data Plane team. In this role, you will be a key technical leader responsible for designing, developing, and delivering high-scale, resilient data plane services that power our "Inference as a Service" offering. You will work at the intersection of distributed systems and specialized AI hardware to ensure our customers can deploy and scale their models with industry-leading performance and reliability. This is a hands-on role, requiring you to be able to develop high quality software while availing of all the productivity boosts granted by the latest AI coding agents.
What You'll Do:
  • Technical Leadership: Act as a technical leader on the team, driving the end-to-end design, development, and delivery of critical data plane components hosting large generative AI models.
  • System Design: Architect and refine system design proposals for our high-scale, multi-tenant AI inference cloud ecosystem, ensuring they meet rigorous availability and resiliency standards.
  • Performance Optimization: Implement and optimize distributed inference hosting using techniques like tensor/data parallelism, KV cache optimizations, and smart routing.
  • Collaboration: Work cross-functionally with Product Managers, customer-facing teams, and other engineering teams to align technical roadmaps with customer needs.
  • Distributed Serving at Scale: Build on Kubernetes-native distributed inference frameworks like llm-d (or alternatives such as NVIDIA Dynamo, Ray Serve, KServe) to deliver prefill/decode disaggregation, KV-cache-aware routing, tiered prefix caching, and wide expert parallelism for MoE models.
  • Flow Control & Load Balancing: Solve the distributed-systems problems unique to LLM serving - inference-aware load balancing on queue depth, cache locality, and predicted latency; flow control and fairness across tenants; autoscaling inference pools; and moving gigabytes of KV-cache between prefill and decode instances with negligible overhead.
  • Open Source Contributions: Contribute upstream to llm-d, vLLM, and the inference gateway ecosystem, and represent DigitalOcean in these communities.
  • Mentorship: Coach and mentor junior engineers, fostering a culture of technical excellence and continuous improvement.
  • Operational Excellence: Maintain and operate critical, high-scale services, utilizing observability tools and defining SLOs to ensure superior platform health.
What You'll Bring to DigitalOcean:
  • AI/ML Domain Knowledge: Hands-on experience hosting large language or multimodal models using inference engines like vLLM, SGLang, or TensorRT.
  • Inference Frameworks: Familiarity with distributed inference serving frameworks such as llm-d, NVIDIA Dynamo, or Ray Serve.
  • Inference Engine Depth: Hands-on experience with vLLM or alternatives (SGLang, TensorRT-LLM, TGI, Modular MAX), including internals like continuous batching, paged attention, and prefix caching.
  • Distributed Inference Fluency: Understanding of why cluster-scale serving is hard: KV-cache locality is partitioned across workers, naive round-robin routing destroys cache hit rates and tail latency, and disaggregated prefill/decode requires fast cross-pod KV transfer (e.g., NIXL).
  • Upstream Track Record: Merged contributions to vLLM, llm-d, SGLang, or similar projects strongly preferred.
  • Architecture Proficiency: Knowledge of common LLM architectures and optimization techniques (e.g., continuous batching, quantization).
  • Software Engineering: Expert-level proficiency in GoLang or Python and familiarity with gRPC.
  • Cloud Operations: Proven experience shipping customer-facing software products and running critical services in a high-scale environment similar to DigitalOcean.
  • Open Source Mindset: Experience integrating and building with open-source software.
Compensation Range:
  • $139,200 - $174,000

*This is a remote role



#LI-Remote

Similar Jobs

More Jobs at DigitalOcean

More Enterprise Technology Jobs

Find similar Senior Software Engineer I - AI Inference Data Plane jobs: