Nutanix

Staff Engineer - Inference /AI

Nutanix$171K — $257K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of experience in software development within a product organization.
  • Strong fundamentals in computer science including data structures and distributed systems.
  • Hands-on expertise in Docker, Kubernetes, and cloud-native architectures.
  • Experience with backend systems in languages like Go, Python, C++, or Rust.
  • Proficient in managing CI/CD pipelines and release automation.
  • Solid understanding of datacenter architecture and virtualization.
  • Experience in designing and deploying software across on-premises and cloud environments.
  • Familiarity with LLM serving concepts and modern AI frameworks.

Responsibilities

  • Architect and design scalable, fault-tolerant services on Kubernetes.
  • Build and operate high-performance services for AI applications.
  • Optimize components of distributed systems and low-level infrastructure.
  • Enhance multi-tenant platform services for diverse AI deployments.
  • Design observability architectures using monitoring technologies.
  • Debug production issues and improve platform reliability.
  • Develop CI/CD pipelines to streamline service delivery.
  • Collaborate with global teams on product development and enhancement.

Benefits

  • RRSP with dollar-for-dollar matching up to 7% of base salary.
  • Comprehensive mental health coverage and paramedical benefits.
  • Fully paid maternity and parental leave along with generous bereavement leave.
  • RSUs and an Employee Stock Purchase Plan at a 15% discount.
  • Hybrid work arrangement promoting in-person collaboration with team alignment.
Full Job Description
Your Role
  • Architect, design, and develop horizontally scalable, containerized, fault-tolerant services on Kubernetes for enterprise AI and LLM workloads.
  • Build and operate high-performance inference and platform services that deliver low-latency, high-throughput experiences for Generative AI and Agentic AI applications.
  • Design and optimize critical system components across the stack, including distributed systems, storage, networking, and low-level infrastructure layers.
  • Develop and enhance multi-tenant platform services supporting on-premises, hybrid, and cloud-based AI deployments.
  • Design and implement scalable observability architectures using technologies such as Prometheus, Grafana, Datadog, OpenTelemetry, and related cloud-native monitoring frameworks.
  • Debug complex production issues, perform root-cause analysis, and improve reliability, resiliency, and operational efficiency of platform services.
  • Build and maintain CI/CD pipelines and deployment automation to accelerate delivery of production-grade services.
  • Design and implement foundational LLM serving capabilities including request routing, rate limiting, token streaming, load balancing, quota management, and usage budgeting.
  • Collaborate closely with globally distributed product management, AI, and software engineering teams to deliver high-quality products in a fast-paced environment.
  • Contribute to all stages of the product lifecycle, including architecture, design, development, testing, experimentation, performance analysis, deployment, and operations.
  • Leverage and contribute to relevant open-source cloud-native and AI ecosystem projects.
  • Review code and design documents, provide feedback on product requirements, and champion engineering excellence across the team.
  • Continuously evaluate emerging technologies and help shape the technical direction of Nutanix's Enterprise AI Platform.

What You Will Bring

Required Qualifications
  • 8+ years of experience developing maintainable, modular, resilient, fail-safe, and long-lived software products within a product development organization.
  • Strong computer science fundamentals including data structures, algorithms, operating systems, networking, and distributed systems.
  • Hands-on experience with Docker, Kubernetes, and cloud-native architectures.
  • Production experience developing backend systems using Go, Python, C++, or Rust.
  • Experience building, owning, and maintaining CI/CD pipelines and release automation end-to-end.
  • Strong understanding of datacenter architecture including compute, storage, networking, and virtualization.
  • Experience designing and deploying software across on-premises, cloud, and hybrid environments.
  • Demonstrated experience designing and tuning high-performance, performance-sensitive system software.
  • Solid understanding of distributed computing, distributed data stores, and large-scale service architectures.
  • Experience diagnosing and resolving production performance issues using observability and monitoring platforms such as Prometheus, Grafana, Datadog, Open Telemetry, or similar tools.
  • Familiarity with LLM serving concepts including rate limiting, token streaming, request scheduling, load balancing, quota management, and usage budgeting.
  • Familiarity with modern LLM concepts including reasoning workflows, tool calling, prompt templates, and agent.
  • Experience building multi-tenant services running on virtualized or containerized infrastructure.
  • Strong communication, collaboration, and problem-solving skills with the ability to work effectively across globally distributed teams.
  • Master's degree in Computer Science or equivalent practical experience.

Bonus Points If You Have Experience With
  • Machine learning frameworks such as PyTorch or TensorFlow.
  • GPU-based systems and acceleration technologies.
  • Modern model-serving platforms such as vLLM, DeepSpeed, Hugging Face TGI, or Triton.
  • Retrieval-Augmented Generation (RAG), vector databases, and AI orchestration frameworks.
  • Open-source contributions or experience working in large distributed codebases.
  • Production AI platforms, LLM APIs, agentic systems, or inference infrastructure.
  • Building or scaling production LLM APIs, including streaming via SSE/WebSockets, prompt guardrails, rate limiting, and usage budgeting.

Learn More About the Technology:https://www.nutanixbible.com/ [nutanixbible.com]

#NAI

Highlighted Benefits (Vancouver, Canada)
Retirement: RRSP with dollar-for-dollar matching up to 7% of base salary
Mental Health: Dedicated mental health coverage plus top-tier paramedical benefits
Family: Fully paid maternity and parental leave and generous bereavement leave, including time for the loss of a pet
Equity: RSUs and Employee Stock Purchase Plan at a 15% discount

Work Arrangement Hybrid: This role operates in a hybrid capacity, blending the benefits of remote work with the advantages of in-person collaboration. In locations where our workplace policy applies (i.e. San Jose, Durham, Mexico City, Vancouver, Bangalore, Pune, Hoofddorp, Belgrade, Barcelona, Singapore, Sydney and Tokyo), employees are expected to work onsite a minimum of 3 days per week to foster collaboration, team alignment, and access to in-office resources. Workplace type may vary based on location and team requirements. Please speak with your recruiter for details. Additional team-specific guidance and norms will be provided by your manager.

Pay Transparency - Role Location The pay range for this position at commencement of employment is expected to be between CAD $171,000 and CAD $257,000 per annual.
However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. The total compensation package for this position may also include other elements, including a sign-on bonus, restricted stock units, and discretionary awards in addition to a full range of medical, financial and/or other benefits (including 401(k) eligibility and various paid time off benefits, such as vacation, sick time, and parental leave), dependent on the position offered. Details of participation in these benefit plans will be provided if an employee receives an offer of employment.

If hired, employee will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors. Our application deadline is 40 days from the date of posting. In good faith, the posting may be removed prior to this date if the position is filled or extended in good faith.

About Nutanix

Nutanix is a cloud computing software company that sells what it calls hyper-converged infrastructure (HCI) appliances and software-defined storage. Nutanix's software delivers a full infrastructure stack that integrates compute, virtualization, storage, networking, and security to power any application, at any scale. Nutanix software runs across different cloud environments to harmonize IT operations and bring frictionless mobility to all applications. Nutanix solutions are used by companies in a wide range of industries, including healthcare, manufacturing, education, and financial services.
Learn more about Nutanix
Size
6,080 employees
Market Cap
$6.3 billion
Industry
Net Income
-$978.4 million
Founded
2009
5 Year Trend
+13.3%
Revenue
$1.3 billion
NASDAQ

Similar Jobs

More Jobs at Nutanix

More Enterprise Technology Jobs

Find similar Staff Engineer - Inference /AI jobs: