Technical Product Manager, AI Inference & Software

Positron AI

$200K — $350K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience in ML systems or inference engineering with ownership of performance trade-offs.
  • Deep knowledge of transformer internals and their hardware implications.
  • Hands-on experience with large-scale production inference serving.
  • Fluency in the open-source inference ecosystem and model adoption practices.
  • Strong skills in performance analysis and benchmarking across systems.
  • Experience in competitive landscape analysis in AI inference and serving stacks.
  • Excellent communication skills, capable of engaging both technical and market-oriented stakeholders.
  • Proficient in using agentic AI in technical planning and requirements drafting.

Responsibilities

  • Lead AI inference and software technical product planning at Positron.
  • Write detailed requirements for the inference software stack.
  • Collaborate with engineering and GTM teams to develop a clear product roadmap.
  • Transform model innovations into concrete engineering requests.
  • Monitor and track advancements in models and inference systems proactively.
  • Gather customer insights from GTM teams to inform the roadmap.
  • Define and manage product lifecycle and documentation processes.

Benefits

  • Comprehensive health care coverage for you and your dependents, including medical, dental, and vision.
  • Unlimited paid time off to promote rest and recharge.
  • Remote-first work culture with equipment provided for home office.
  • Competitive salary plus equity opportunities.
  • 401(k) with company matching from day one.
  • Life and disability insurance, with optional enhancements.
Full Job Description
Role Overview

Positron AI is looking for a Technical Product Manager to own AI inference and software technical product planning end to end. In this role, you will be the person who translates where models and inference systems are heading into concrete, well-scoped requirements for our inference software stack, spanning model coverage, numerics, inference-engine features and modes, serving-stack capabilities, and our managed service.

This is a deeply technical planning role that sits at the intersection of engineering, go-to-market, and the broader inference ecosystem. You will track the model frontier as a discipline, convert that movement into engineering requests before it becomes a customer escalation, and serve as the connective tissue between our engineering organization, our GTM teams, and our ecosystem partners. You will also be expected to use agentic AI daily as a core part of how the planning function operates.

Key Responsibilities

Product Requirements and Roadmap

  • Serve as the leader at Positron for all aspects of AI inference and software technical product planning.
  • Write requirements for the inference software stack spanning model coverage, numerics, inference-engine features and modes, serving-stack features and modes, and managed-service capabilities.
  • Partner with engineering and GTM to build and communicate a clear, defensible roadmap.
  • Create scope and feasibility frameworks that convert model and inference-system innovations into tangible engineering requests.

Market and Frontier Tracking

  • Keep planning ahead of where models and inference systems are going, tracking the model frontier as a discipline and converting movement into requirements before it surfaces as a customer escalation.
  • Work closely with GTM teams to understand customer and market needs and feed them back into the roadmap.
  • Create competitive briefings covering inference providers, serving stacks, and adjacent hardware platforms.
  • Engage key ecosystem partners, including model labs, open-source runtimes, and serving and orchestration partners, to understand their technology roadmaps.

Lifecycle, Documentation, and Process

  • Define and manage the software product lifecycle, including versions and release trains, feature modes, model catalog, and deprecation policy.
  • Ensure our software products are well documented for both internal and customer-facing audiences.
  • Streamline and automate the product planning process using agentic AI.

Required Qualifications

  • 10+ years of experience across ML systems, inference infrastructure, or serving-stack engineering, with direct ownership of performance or architecture trade-offs.
  • Deep expertise in transformer internals at the operator level, including attention variants, MoE routing, KV-cache mechanics, and quantization formats along with their hardware implications.
  • Hands-on experience with production inference serving at scale, covering multi-tenancy, latency SLAs such as TTFT and TPOT, batching and scheduling, disaggregated serving, KV-cache management, and observability.
  • Working fluency in the open-source inference ecosystem, including vLLM and SGLang-class runtimes, kernels, model ingestion, and how models are released, quantized, and adopted in practice.
  • Strong performance analysis skills spanning models and systems, including utilization reasoning, tokens per dollar and tokens per watt arithmetic, and benchmark design, with the ability to build and defend the math personally.
  • Demonstrated experience in competitive landscaping and analysis of inference providers and serving stacks, gained at a model lab, an inference API provider, or an AI hardware company.
  • A proven ability to learn quickly and span the full stack, from model-architecture details up to fleet-scale serving systems, while staying current with the model and inference landscape.
  • Excellent communication and interpersonal skills, with comfort navigating uncertainty and driving a process of idea and decision socialization.
  • Confidence being the most technically grounded person in a GTM room and the most market-aware person in an engineering room.
  • A strong instinct for owning decision history, serving as the documented answer to "why did we choose X," including with executive leadership.
  • Daily, hands-on use of agentic AI in real technical work, building and running agent workflows for research, analysis, and requirements drafting, with the judgment to verify and own everything the agents produce.

Preferred Qualifications

  • Prior experience at an AI hardware or custom silicon company, with exposure to the realities of bringing a new accelerator platform to market.
  • Direct contribution to or close engagement with open-source inference runtimes or serving projects.
  • Experience defining and operating a managed inference service, including model catalog and deprecation policy.
  • A track record of building internal automation or agentic workflows that measurably improved a planning or research function.

Leveling & Scope

While this role is currently posted at a specific level, we are a growth-oriented organization and are open to hiring at a more senior level for the right candidate. Please note that this job description serves as a focused but generalized overview of the role; specific responsibilities and impact expectations will be tailored to the experience and seniority of the final hire.

Why Join Us?

  • You will shape the software roadmap for a purpose-built inference accelerator, working at the layer where model architecture, systems performance, and real customer workloads meet.
  • You will have unusually direct influence, defining what gets built across the inference stack and seeing it land in silicon-backed products that compete on performance per dollar and performance per watt.

Compensation & Benefits

The base salary range for this role is $200,000 - $350,000.

Please note that the figures provided represent the base salary range only and do not include other elements of our total compensation package, equity, or comprehensive benefits.

At Positron AI, we value the unique expertise each candidate brings. While the range above reflects our typical expectation for the position, we reserve the flexibility to exceed this range for candidates whose specialized skills, significant experience, or unique qualifications fall outside the standard scope of the role. Final offers are determined based on a variety of factors, including internal equity, and individual impact.

Benefits & Perks

We want you to do your best work and feel confident that you and your family are taken care of. That means comprehensive coverage, real time to rest, and support for your future.

Health and wellness

  • Fully company-paid medical, dental, and vision insurance for you and your dependents
  • Company-paid life and disability coverage, with voluntary options to add more
  • Supplemental hospital, critical illness, and accident coverage available

Time off and flexibility

  • Unlimited paid time off, we encourage everyone to truly unplug and recharge
  • 13 paid company holidays
  • Remote-first culture with a company-provided computer and home office setup

Compensation and future

  • Competitive salary and equity
  • 401(k) with company matching, eligible from day one

Visa Support

This position is open to candidates currently authorized to work in the U.S. We cannot provide new visa sponsorship for this role but are open to facilitating H-1B visa transfers for eligible candidates.

Similar Jobs

More Jobs at Positron AI

More Information Technology Jobs

Find similar Technical Product Manager, AI Inference & Software jobs: