About the roleWEKA is looking for a Product Manager to own the roadmap and go-to-market for Augmented Memory Grid (AMG), part of the NeuralMesh platform. This is a deeply technical PM role sitting at the intersection of AI inference infrastructure, high-performance networking, and enterprise storage. You will work directly with engineering, GPU/inference partners (NVIDIA, hyperscalers, GPU clouds), and enterprise customers running large-scale LLM inference to define what AMG needs to do next.
Bring Your Expertise - and Your Passion- Leadership Skills: Strong leadership skills with a history of successfully leading cross-functional teams. Product Managers are expected to inspire and motivate team members to achieve ambitious goals while maintaining a collaborative and positive working environment. You understand how to influence without authority, and your recall of meaningful details supports verbal and written agility.
- Strategic Vision: You are a strategic thinker who can develop and execute product strategies that align with market trends and customer needs, as well as think critically about existing strategies. You have a proven ability to translate strategic goals into actionable plans and deliver results.
- Communication Skills: You have excellent communication and interpersonal skills, with the ability to articulate complex technical concepts to both technical and non-technical stakeholders. You are comfortable presenting product strategies and roadmaps to internal teams and external customers.
What you'll do- Own the AMG product roadmap: KV-cache/prefix-cache offload, memory tiering, and integration with inference engines and orchestration layers (vLLM, NVIDIA Triton/TensorRT-LLM/NIM, Kubernetes-based serving).
- Partner with engineering to define architecture trade-offs across GPU memory, networking (RDMA, GPUDirect, NVMe-oF), and distributed storage - translating inference performance bottlenecks (time-to-first-token, throughput, context length) into product requirements.
- Work directly with enterprise customers and GPU cloud partners: Nebius, CoreWeave, TogetherAI, etc., running production inference workloads to gather requirements, validate benchmarks, and prioritize features that reduce cost-per-token and improve SLAs at scale.
- Partner with NVIDIA and other silicon/inference-stack partners on joint roadmap and certification work.
- Define and track benchmarks (TTFT, throughput, cache hit rate) that demonstrate AMG's value versus standard GPU-memory-only inference.
- Support sales and field teams with technical positioning, competitive differentiation, and enterprise deal support.
Must-have qualifications- Inference ecosystem depth: hands-on product or engineering experience with LLM inference serving - vLLM, NVIDIA Triton/TensorRT-LLM/NIM, Ray Serve, or comparable - and fluency in concepts like KV-cache, prefix/context caching, quantization, and batching strategies.
- Model & systems familiarity: working knowledge of how modern LLMs are served in production (context windows, multi-tenant serving, GPU scheduling) well enough to translate model-level constraints into infrastructure requirements.
- Networking/infrastructure fluency: comfort with the fundamentals of high-performance networking and distributed systems - RDMA, GPUDirect Storage, NVMe-oF, or equivalent - and how they affect inference performance.
- Enterprise customer experience: track record working directly with large enterprise accounts - requirements gathering, production deployments, SLAs - not solely self-serve/PLG products.
- 10+ years of product management experience, ideally with some portion in infrastructure, ML platforms, or developer-facing technical products.
Nice-to-have- Prior experience at a GPU cloud, inference platform, or AI infrastructure startup.
- Familiarity with storage systems (parallel/distributed file systems, object storage) in AI/ML pipelines.
- Experience partnering directly with NVIDIA or other accelerator/silicon vendors.