About the RoleWe're looking for a Forward Deployed Engineer to make Inferact successful inside real customer environments. You'll work directly with customers to deploy, integrate, debug, and optimize vLLM-powered inference systems across cloud, Kubernetes, GPU, networking, and model-serving environments.
This is a hands-on engineering role, not a traditional pre-sales position. You'll move from architecture discussions to implementation, own difficult production problems end-to-end, and work closely with core product and engineering teams to turn what you learn in the field into reusable product capabilities. Your work will directly affect customer time-to-value and how Inferact's platform evolves.
Skills and QualificationsMinimum qualifications:
- Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.
- Strong software engineering ability in Python, Go, TypeScript, or similar, with experience building production-quality integrations, tooling, services, automation, or prototypes.
- Hands-on experience deploying or operating ML systems, model serving, AI infrastructure, cloud platforms, Kubernetes, or high-scale backend systems in production.
- Ability to work directly with sophisticated customer engineering teams, understand ambiguous technical environments, and personally drive implementations and debugging to resolution.
- Strong systems debugging skills across application, runtime, infrastructure, networking, identity, storage, observability, and distributed-system boundaries.
- Ability to reason about latency, throughput, batching, model/runtime compatibility, scaling, reliability, and cost tradeoffs in production inference environments.
- High ownership and strong technical communication, with the judgment to distinguish one-off customer work from problems that should become reusable product capabilities.
Preferred qualifications:
- Experience with vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, BentoML, or other LLM inference and model-serving systems.
- Experience with NVIDIA or AMD GPUs, CUDA / ROCm, GPU scheduling, multi-GPU serving, or accelerator-backed infrastructure.
- Experience deploying infrastructure software into enterprise, regulated, security-sensitive, or bring-your-own-cloud environments.
- Experience building APIs, SDKs, CLIs, developer tooling, deployment platforms, control planes, or infrastructure products used by technical teams.
- Experience profiling latency, throughput, concurrency, GPU utilization, bottlenecks, and performance regressions.
Bonus points if you have:
- Contributed to open-source ML systems, inference infrastructure, cloud infrastructure, Kubernetes, or developer tooling.
- Worked in a forward-deployed, customer engineering, field engineering, or highly technical solutions role where you personally wrote and shipped code.
- Built deployment playbooks, reference architectures, automation, or tooling that materially reduced customer time-to-production.
- Resolved severe customer-facing production issues that crossed multiple technical layers and required close partnership with core engineering.
- Turned repeated customer problems into reusable product features, abstractions, documentation, or platform improvements.
Logistics- Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
- Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
- Visa sponsorship: We sponsor visas on a case-by-case basis.
- Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.