Job Summary
We are seeking a Real-Time Inference Engineering Lead to design, build, and industrialize low-latency, resilient model-serving services for real-time predictive AI use cases. This hands-on engineering leadership role will establish reusable patterns for online inference architecture, deployment, API services, capacity management, observability, reliability, CI/CD, and production operations across public cloud and on-premises environments. The ideal candidate will have strong experience with model serving, Kubernetes, APIs, performance optimization, reliability engineering, and production operations, with the ability to guide application, data science, and ML engineering teams in safely deploying predictive models at enterprise scale.
Key Responsibilities
• Define target architecture and engineering standards for real-time predictive model-serving services across cloud and on-premises environments.
• Design, build, test, deploy, and operate scalable online inference services meeting latency, throughput, availability, resiliency, and security requirements.
• Establish reusable model-serving patterns for synchronous APIs, asynchronous inference, batch-adjacent processing, and event-driven real-time use cases.
• Build standardized model deployment approaches covering model packaging, versioning, release promotion, canary deployment, rollback, and retirement.
• Design secure inference API patterns covering authentication, authorization, traffic management, rate limiting, auditability, and enterprise integration.
• Engineer Kubernetes-based serving platforms using GKE, OpenShift, and related container orchestration technologies.
• Implement autoscaling, resource allocation, quota management, capacity planning, and workload-isolation controls.
• Conduct performance engineering, load testing, stress testing, and failure testing across expected and peak production workloads.
• Identify and implement latency optimizations across model initialization, feature retrieval, network paths, API handling, runtime configuration, and infrastructure utilization.
• Define and implement monitoring, telemetry, dashboards, alerts, SLIs, SLOs, and error-budget practices.
• Partner with ML platform, data engineering, application engineering, security, and operations teams to integrate model services with governed data, feature, network, and identity capabilities.
• Implement CI/CD and automated validation for model-serving services, infrastructure configuration, APIs, performance benchmarks, and release-readiness checks.
• Develop operational runbooks, incident-response procedures, support models, and production-readiness documentation.
• Drive reliability improvements through root-cause analysis, capacity reviews, resiliency testing, disaster-recovery planning, and continuous operational improvement.
• Mentor engineers and establish reusable technical documentation, reference implementations, and knowledge-transfer materials.
Required Qualifications
• 8+ years of experience in software engineering, platform engineering, cloud engineering, SRE, or infrastructure engineering.
• 4+ years of experience designing, building, or operating production APIs, distributed systems, platform services, or real-time data and ML workloads.
• Demonstrated experience leading technical design and engineering delivery for highly available, performance-sensitive production services.
• Strong experience with online inference architecture, model-serving frameworks, or predictive-model deployment patterns.
• Hands-on experience designing and operating RESTful, gRPC, or event-driven APIs.
• 4+ years of strong production experience with Kubernetes and container platforms, including GKE, OpenShift, or comparable environments.
• Experience with autoscaling, resource management, capacity planning, performance testing, and load testing for distributed services.
• Experience implementing observability, monitoring, dashboards, alerts, SLIs, SLOs, and incident-management practices.
• Experience with CI/CD, Git-based development, automated testing, deployment automation, and production-release practices.
• Strong understanding of resiliency, high availability, fault tolerance, disaster recovery, and operational support for critical services.
• Ability to collaborate effectively with data science, ML engineering, platform engineering, application, security, and business teams.
Required Skills / Knowledge
• Online inference and low-latency model-serving architecture.
• Model deployment, versioning, routing, rollout, rollback, and lifecycle management.
• REST APIs, gRPC, API gateways, authentication, authorization, traffic management, and API observability.
• Kubernetes, GKE, OpenShift, containers, service meshes, ingress, workload scheduling, and autoscaling.
• Performance engineering, load testing, stress testing, benchmarking, profiling, and latency optimization.
• Monitoring, telemetry, distributed tracing, dashboards, alerting, SLIs, SLOs, and error budgets.
• CI/CD, automated testing, deployment automation, infrastructure-as-code, and release controls.
• Resiliency engineering, high availability, capacity controls, incident response, root-cause analysis, and operational runbooks.
• Cloud and on-premises platform operations, networking, identity, data protection, and secure production delivery.
Preferred Qualifications
• Experience with Vertex AI endpoints, KServe, Seldon, NVIDIA Triton Inference Server, MLflow deployments, or comparable model-serving technologies.
• Experience deploying and operating models on GCP, Azure, AWS, private cloud, or hybrid-cloud environments.
• Experience with service mesh, API gateway, traffic-routing, or edge-serving technologies.
• Experience serving high-volume, customer-facing, fraud, risk, personalization, decisioning, or other latency-sensitive predictive models.
• Experience with feature serving, online feature stores, caching, streaming platforms, or real-time data enrichment.
• Experience with Terraform, Helm, Argo CD, Jenkins, GitHub Actions, GitLab CI, or similar automation tooling.
• Experience in banking, financial services, healthcare, insurance, or other regulated enterprise environments.
• Experience participating in a 24x7 operational support model for high-priority production services.
Expected Outcomes
• Standardized, production-ready real-time inference architecture and reusable model-serving patterns.
• Reliable online inference services meeting defined latency, throughput, availability, and resiliency objectives.
• Automated deployment, testing, monitoring, capacity management, and rollback capabilities for predictive models.
• Clear operational dashboards, SLOs, alerts, runbooks, and readiness evidence for real-time services.
• Improved engineering productivity and faster adoption of secure, scalable real-time predictive AI capabilities across the Cortex portfolio.