Technical Product Manager, Infrastructure

Vast.ai Inc

$170K — $240K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years as a backend engineer and 3+ years in product management
  • Hands-on experience with security, scalability, reliability, infrastructure tooling, distributed systems, observability, and compliance
  • Industry experience in software infrastructure, developer tooling, AI/ML, cloud computing, or two-sided marketplaces
  • Experience at a fast-paced startup or rapid-growth team

Responsibilities

  • Drive the infrastructure, security, and reliability roadmap
  • Sequence competing priorities and translate requirements into actionable specs for engineering
  • Own metrics related to platform uptime, fleet reliability, and cost-to-serve
  • Focus on scalability and performance to keep core systems fast
  • Enhance platform security and comply with enterprise requirements
  • Improve observability and tooling for engineers and hosts

Benefits

  • Comprehensive health, dental, vision, and life insurance
  • 401(k) with company match
  • Meaningful early-stage equity
  • Onsite meals, snacks, and close collaboration with founders
  • Ambitious, fast-paced startup culture where initiative is rewarded
  • Ample AI agent budget
Full Job Description
About the Role

Vast.ai is seeking a Technical Product Manager to drive the backend our GPU cloud marketplace runs on. This is the software behind every live GPU rental (over 700k transactions a month): the daemon on every host's machines, the orchestration that powers each instance, and the infrastructure and test systems that let us ship it all at high velocity.

The Scale: With over 20k GPUs, our AI cloud platform powers thousands of bleeding-edge training runs and critical production workloads for 120k+ developers all over the planet. This is a product with real scale, real data, and real users from day one. The work you ship moves tangible revenue in weeks, not quarters.

The Challenge: We don't own the GPUs. Thousands of independent hosts price and operate them, and no two machines are alike. Your job is to make that heterogeneous, decentralized supply behave like the top-tier cloud our customers expect: fast, secure, and reliable.

What You'll Own

You'll drive the infrastructure, security, and reliability roadmap: sequencing competing priorities, turning non-functional requirements into specs engineering can build, and seeing them through until they ship. You'll own the metrics that prove it worked: platform uptime, fleet reliability, cost-to-serve. Initial focus areas:
  • Scalability & performance. Keep the core systems fast and ahead of demand as the marketplace grows, from database performance to end-to-end latency.
  • Security, trust & compliance. Harden the platform against attacks and abuse, and build the compliance roadmap enterprise customers need.
  • Observability & infra tooling. Give every engineer a clear view of platform health, and every host a clear view of their fleet. Own the metrics, tracing, and tooling behind it.


Ideal Experience
  • 3+ years as a backend engineer AND 3+ years in product management
  • Hands-on experience with several of: security, scalability, reliability at scale, infrastructure tooling, distributed systems, observability, abuse/fraud prevention, compliance
  • Industry experience in one or more of the following areas: software infrastructure, developer tooling (APIs, SDKs, CLIs), AI/ML, cloud computing, GPUs, or two-sided marketplaces
  • Experience at a fast-paced startup or rapid-growth team


You Are
  • AI-native. You work with AI agents every day, and you want to build the compute layer the AI era runs on.
  • Deeply technical. You came up as an engineer, and you reason about scalability, security, and efficiency as product concerns, not afterthoughts.
  • Persuasive across the stack. You can make a technical tradeoff legible to c-suite and a thankless migration compelling to the team doing it.
  • Biased toward action. You'd rather ship the fix today than present the plan next week.


Tech Stack

Python, C++, SQL/PostgreSQL, Redis, Linux, Docker, AWS, Terraform, REST APIs, LLMs/AI agents

Interview Process
  • Initial screening (virtual,)
  • Deep dive on the role and your experience (virtual)
  • Live product and technical assessment (virtual)
  • In person panel interview + product and technical assessment (on-site)


Annual Salary Range

$170,000-$240,000 base depending on experience, plus equity. We're profitable, venture-backed, and growing, so the equity is a stake in a real company with a real valuation, not a lottery ticket.

Benefits
  • Comprehensive health, dental, vision, and life insurance
  • 401(k) with company match
  • Meaningful early-stage equity
  • Onsite meals, snacks, and close collaboration with founders/tech leaders
  • Ambitious, fast-paced startup culture where initiative is rewarded
  • Ample AI agent budget


LOCATION: On-site at our office in Los Angeles (Westwood).

Similar Jobs

More Enterprise Technology Jobs

Find similar Technical Product Manager, Infrastructure jobs: