Staff AI Runtime Engineer

FlexAI

• $180K — $225K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of experience in systems/software engineering with a focus on AI runtime and distributed systems.
  • Experience delivering PaaS services.
  • Proven skills optimizing and scaling deep learning runtimes like PyTorch or TensorFlow.
  • Strong programming proficiency in Python and C++ (knowledge of Go or Rust is a bonus).
  • Familiarity with distributed training frameworks and low-level performance tuning.
  • Experience with multi-GPU, multi-node, or cloud-native AI workloads.
  • Solid understanding of containerized workloads and production failure recovery.

Responsibilities

  • Lead the design and development of the core AI runtime architecture.
  • Create resilient runtime features for our PyTorch stack, including dynamic node scaling.
  • Optimize reliability and performance of distributed training and inference workflows.
  • Enhance system performance across AI training and inference pipelines.
  • Develop libraries supporting the model lifecycle from training to deployment.
  • Implement observability and diagnostics for deep learning workloads.
  • Mentor junior engineers and guide technical discussions across teams.

Benefits

  • Competitive salary and benefits package.
  • Opportunity to work on cutting-edge AI infrastructure.
  • Develop products utilized by both developers and enterprises.
  • High ownership and fast execution to drive real impact.
  • Collaborative environment with a high-caliber team.
Full Job Description
Role Overview

At FlexAI, we're building a high-performance, cloud-agnostic AI compute platform designed for next-generation training and inference workloads. As a Staff AI Runtime Engineer, you'll play a pivotal role in the design, development, and optimization of the core runtime infrastructure that powers distributed training and deployment of large AI models (LLMs and beyond).

This is a hands-on leadership role - perfect for a systems-minded software engineer who thrives at the intersection of AI workloads, runtimes, and performance-critical infrastructure. You'll own critical components of our PyTorch-based stack, lead technical direction, and collaborate across engineering, research, and product to push the boundaries of elastic, fault-tolerant, high-performance model execution.

What You'll Do

Lead Runtime Design & Development:
  • Own the core runtime architecture supporting AI training and inference at scale.
  • Design resilient and elastic runtime features (e.g. dynamic node scaling, job recovery) within our custom PyTorch stack.
  • Optimize distributed training reliability, orchestration, and job-level fault tolerance.

Drive Performance at Scale:

  • Profile and enhance low-level system performance across training and inference pipelines.
  • Improve packaging, deployment, and integration of customer models in production environments.
  • Ensure consistent throughput, latency, and reliability metrics across multi-node, multi-GPU setups.

Build Internal Tooling & Frameworks:

  • Design and maintain libraries and services that support model lifecycle: training, checkpointing, fault recovery, packaging, and deployment.
  • Implement observability hooks, diagnostics, and resilience mechanisms for deep learning workloads.
  • Champion best practices in CI/CD, testing, and software quality across the AI Runtime stack.

Collaborate & Mentor:

  • Work cross-functionally with Research, Infrastructure, and Product teams to align runtime development with customer and platform needs.
  • Guide technical discussions, mentor junior engineers, and help scale the AI Runtime team's capabilities.


What You'll Need to Be Successful

  • 8+ years of experience in systems/software engineering, with deep exposure to AI runtime, distributed systems, or compiler/runtime interaction.
  • Experience in delivering PaaS services.
  • Proven experience optimizing and scaling deep learning runtimes (e.g. PyTorch, TensorFlow, JAX) for large-scale training and/or inference.
  • Strong programming skills in Python and C++ (Go or Rust is a plus).
  • Familiarity with distributed training frameworks, low-level performance tuning, and resource orchestration.
  • Experience working with multi-GPU, multi-node, or cloud-native AI workloads.
  • Solid understanding of containerized workloads, job scheduling, and failure recovery in production environments.

Nice to Have

  • Contributions to PyTorch internals or open-source DL infrastructure projects.
  • Familiarity with LLM training pipelines, checkpointing, or elastic training orchestration.
  • Experience with Kubernetes, Ray, TorchElastic, or custom AI job orchestrators.
  • Background in systems research, compilers, or runtime architecture for HPC or ML.
  • Start up previous experience

This position is In-Person and located at our Santa Clara, CA Office.

What We Offer

  • A competitive salary and benefits package
  • Work on cutting-edge AI infrastructure
  • Build products used by developers and enterprises
  • High ownership, fast execution, real impact
  • Collaborative, high-caliber team


The pay range for this role is:

180,000 - 225,000 USD per year (US)

Similar Jobs

More Information Technology Jobs

Find similar Staff AI Runtime Engineer jobs: