Key Responsibilities:
AI Architecture & Technical Leadership
- Define and lead the technical architecture for enterprise-scale AI and ML platforms.
- Design scalable, resilient, and reusable AI systems capable of supporting mission-critical workloads.
- Establish architectural standards, engineering patterns, and best practices for AI deployment and operations.
- Drive technical decisions around model serving, inference optimization, agent architectures, orchestration frameworks, observability, and AI infrastructure.
Productize AI Research
- Partner closely with AI researchers to transform cutting-edge prototypes into production-grade solutions.
- Lead efforts to operationalize advanced AI capabilities across areas such as:
- Large Language Models (LLMs)
- Trustworthy and Responsible AI
- Establish repeatable pathways that accelerate innovation-to-production cycles.
- Ensure production solutions maintain scientific rigor while meeting enterprise engineering standards.
- Bridge the gap between research breakthroughs and sustainable business value.
Engineering Excellence & Scalability
- Solve the organization's most complex AI engineering and scalability challenges.
- Design systems that operate reliably at enterprise scale while balancing performance, latency, governance, security, and cost.
- Drive adoption of MLOps, LLMOps, and AI platform engineering best practices.
- Improve the robustness, maintainability, observability, and operational readiness of our AI products.
- Identify and eliminate architectural bottlenecks that impact scale, reliability, or client experience.
- Raise standards through coaching, architecture reviews, design guidance, and technical leadership.
Production Reliability & Operational Leadership
- Own the operational excellence, reliability, performance and availability of our products.
- Lead technical response and resolution efforts for complex production incidents, performance degradation, model failures, and system outages.
- Serve as the senior technical escalation point for the team's most challenging production challenges.
- Establish best practices for AI system monitoring, observability, alerting, incident management, capacity planning, and service-level objectives (SLOs).
- Mentor and lead junior engineers in troubleshooting, root cause analysis, operational decision-making, and incident response.
- Drive post-incident reviews focused on learning, continuous improvement, and long-term corrective actions.
- Develop operational processes that ensure AI solutions remain secure, scalable, performant, and reliable for business-critical use cases.
- Partner with product, infrastructure, security, and support teams to proactively identify operational risks and continuously improve service reliability.
Mentorship & Thought Leadership
- Mentor AI and ML engineers within the team.
- Foster a culture of technical excellence and operational ownership where engineers are accountable not only for building systems, but also for running and supporting them successfully in production.
- Represent our team as a thought leader in scalable AI deployment, operational excellence, and responsible AI practices.
Required Qualifications
- 10+ years of experience in software engineering, machine learning engineering, AI engineering, or related technical disciplines.
- Deep expertise designing, deploying, and supporting large-scale AI and ML systems in production environments.
- Demonstrated success leading complex technical initiatives from concept through deployment and ongoing operations.
- Strong knowledge of software architecture, reliability engineering, observability, ML Ops, DevOps, and cloud technologies.
- Proven ability to mentor engineers and lead teams through highly complex technical and operational challenges.
Preferred Qualifications
- Experience with foundation models, Large Language Models, and agentic AI architectures.
- Experience deploying agentic AI systems and multi-agent workflows.
- Experience with Trustworthy AI, Responsible AI, AI governance, or model risk management frameworks.
- Experience optimizing large-scale inference systems and AI infrastructure.
- Experience working in highly regulated environments and mission-critical production systems.
What You'll Gain
This role offers a unique opportunity to operate at the forefront of applied artificial intelligence and help bridge world-class research with real-world impact.
You will:
- Work directly with world-class AI researchers on breakthrough technologies and next-generation AI capabilities.
- Own a critical position in the pipeline that transforms cutting-edge research into client value.
- Tackle some of the most difficult AI engineering, scalability, and operational challenges in the industry.
- Build AI capabilities that deliver meaningful business outcomes for clients.
- Develop deep expertise in operating advanced AI systems at scale while collaborating with leaders across research, product, and engineering.
Special FactorsSponsorshipVanguard is not offering visa sponsorship for this position.