Where this role sitsThis role owns how Sapience AI serves intelligence at scale. You build and run the inference infrastructure that turns the models and reasoning behind MINERVA and COGENT into fast, reliable, affordable answers for members.
You work where models meet production: serving, scaling, latency, and cost. Every time a member asks the platform something, your systems are what make the answer arrive quickly and hold up under load.
You sit close to applied AI, research, and platform engineering, and you make the difference between a model that works in a notebook and one that serves a community reliably.
Why this role existsCollective intelligence is only useful if it is fast and dependable. Members will not wait, and communities cannot rely on a platform that buckles under load or costs too much to run.
Serving modern AI at scale is a hard systems problem: large models, tight latency budgets, expensive hardware, and demand that spikes. It takes real infrastructure engineering to get right.
The AI/ML Infrastructure and Systems Engineer owns that problem. You make inference fast, reliable, and affordable, so the intelligence the platform promises actually reaches members.
What you will own (Areas of Responsibility)You hold seven areas of responsibility across inference infrastructure. Each one is yours to set direction on, build, and measure.
1. Inference serving and systems- Build and operate the inference systems that serve models and reasoning behind MINERVA and COGENT.
- Own the serving path end to end, from request to response, under real load.
- Make serving robust to failure, traffic spikes, and change.
2. Latency and performance- Drive down latency so members get fast answers, and keep it low as the platform grows.
- Optimize the full path, including model execution, batching, caching, and retrieval.
- Profile relentlessly and remove the bottlenecks that matter.
3. Scale and reliability- Scale inference to more members, more communities, and heavier reasoning without losing reliability.
- Build autoscaling, load management, and graceful degradation.
- Own the reliability of the serving layer as a first-order responsibility.
4. Cost and efficiency- Own the economics of inference, including GPU and hardware efficiency at scale.
- Improve utilization and cut waste without hurting quality or speed.
- Make cost a designed property, not a surprise.
5. Model deployment and lifecycle- Build the paths that take models and reasoning components from research to production safely.
- Support rollout, versioning, and rollback with confidence.
- Give applied AI and research a fast, safe route to ship.
6. Observability and operations- Instrument the serving layer so it can be measured, debugged, and trusted.
- Build the alerts, dashboards, and tooling that keep inference healthy.
- Reduce the operational burden of running AI at scale.
7. Hardware, accelerators, and platform choices- Make sound choices about accelerators, runtimes, and serving frameworks.
- Balance performance, cost, and maintainability in platform decisions.
- Keep the stack current as inference technology moves.
AI-augmented ways of workingYou build the infrastructure that serves AI, and you use AI in building it, to generate tooling, reason about performance, and move faster, while you own correctness, reliability, and cost.
The standard is human in partnership: AI accelerates the work, you own the judgment, the interpretation, and the call. The people who create the most value here are not the ones producing the most output. They are the ones turning evidence into clear, durable decisions.
What this role is notTo keep the boundary clear:
- This is not a model research role. You serve and scale models; you do not develop new model architectures.
- This is not a general backend role. Your center of gravity is inference systems, performance, and hardware efficiency.
- This is not a data engineering role. You partner with data teams, but you own serving, not the data platform.
- This is not a best-effort prototype role. You are accountable for production inference members depend on.
What success looks likeWe measure this role on outcomes the team can see:
- Fast answers. Members get low-latency responses, and latency stays low as the platform grows.
- Reliable at scale. Inference holds up under load, spikes, and growth.
- Affordable intelligence. Inference cost per answer improves as usage rises.
- Safe rollouts. Models and reasoning components ship, version, and roll back with confidence.
- Operable serving. The serving layer is instrumented, debuggable, and healthy.
- Sound platform choices. Accelerator and framework decisions age well.
Who you areRequired qualifications- Five or more years in infrastructure, systems, or ML infrastructure engineering.
- Hands-on experience serving ML or LLM models in production at scale.
- Deep understanding of latency, throughput, and performance optimization.
- Experience with GPUs or accelerators and their efficient use.
- Strong systems programming and distributed systems fundamentals.
- A track record of reliable, cost-aware production systems.
- Fluency with observability and operational excellence.
Preferred qualifications- Experience with inference-serving frameworks and model runtimes.
- Experience optimizing LLM inference, including batching, quantization, and caching.
- Familiarity with retrieval systems and their performance characteristics.
- Experience owning cost and capacity for AI workloads.
- Exposure to neuro-symbolic or agentic systems in production.
How you work- You name the real bottleneck before reaching for a fix.
- You measure before and after, and you trust evidence over intuition.
- You treat reliability and cost as first-order, alongside speed.
- You build systems others can operate.
- You share tooling and knowledge across the team.
Skills & Competencies- Inference serving architecture and optimization.
- Latency, throughput, and performance engineering.
- GPU and accelerator efficiency.
- Distributed systems, scaling, and reliability engineering.
- Model deployment, versioning, and rollout.
- Observability and production operations for AI.
- Cost and capacity management for inference.
Services & Tools Experience- Inference-serving frameworks and model runtimes (for example vLLM, TensorRT, Triton-class systems).
- GPU tooling, CUDA-class ecosystems, and accelerator runtimes.
- Kubernetes, containers, and cloud platforms (AWS, GCP, or Azure).
- Observability stacks (metrics, tracing, and logging).
- Python and a systems language such as Go, Rust, or C++.
- Caching, queuing, and load-management systems.
- Serving the models and reasoning behind the MINERVA platform and COGENT architecture.
Prior Experience & Background- Prior ML infrastructure, platform, or systems engineering at a software or AI company.
- Experience serving models in production under real load.
- A track record of performance and reliability improvements at scale.
- Experience owning inference cost is a plus.
Cross-functional partnersYou work most closely with Applied AI, Neuro-Symbolic AI, Platform Engineering, and Research. You serve the models and reasoning that power MINERVA and the COGENT architecture, and you own how they run in production.
How we hireWe review every application, and we encourage you to apply even if you do not match every line above. Research shows that talented people, especially those from underrepresented communities, often hold back when they do not meet every qualification. If that is the only thing holding you back, apply anyway.
CompensationBase Salary: $204,000 - $216,000 + early stage equity
Generous health and wellness benefits