Job DescriptionThe Senior Manager will lead a team of engineers while serving as a hands-on technical leader responsible for defining, building, and operating next-generation AI systems on Oracle Cloud Infrastructure (OCI). This leader will set the architecture and engineering direction for production-grade agentic AI platforms, autonomous workflows, scalable inference infrastructure, and enterprise AI applications used in large-scale, business-critical environments.
This role requires an experienced engineering manager who can build and develop high-performing teams, translate ambiguous product and platform goals into a durable technical strategy, and drive execution across multiple organizations. The successful candidate will be accountable for hiring, mentoring, performance management, technical planning, and delivery, while remaining actively involved in system design, prototyping, coding, code reviews, operational readiness, and incident follow-up.
The ideal candidate combines deep distributed-systems expertise with practical, hands-on experience building AI agents and AI-native applications. This includes developing and orchestrating LLM-based agents, tools, APIs, memory systems, retrieval pipelines, evaluations, guardrails, and cloud-service integrations. The candidate should be comfortable writing production code, debugging complex systems, and guiding engineers through difficult architectural and implementation decisions.
The expectation is to lead the team in shipping, scaling, and operating reliable, secure, observable, and cost-efficient AI systems, while raising both the engineering and management bar across the organization. This leader will establish strong execution practices, promote operational excellence, and ensure the team delivers measurable business and customer outcomes.
ResponsibilitiesResponsibilities
- Lead and develop a team responsible for OCI AI platform capabilities, including agent execution, inference, orchestration, evaluation, and observability.
- Set the technical direction for production-grade agentic AI systems that support reasoning, planning, tool use, multi-step workflows, and human escalation.
- Remain hands-on in architecture, prototyping, coding, debugging, and code reviews for critical AI-agent components.
- Guide the development of services for tool calling, memory, context management, MCP integration, retrieval, multi-agent coordination, policy enforcement, and evaluation.
- Own delivery across distributed systems optimized for reliability, performance, security, cost, and multi-tenant operation.
- Translate broad goals into roadmaps, staffing plans, milestones, and measurable outcomes.
- Partner across infrastructure, security, data, product, and application teams to drive execution.
- Establish AgentOps and LLMOps practices for tracing, monitoring, testing, safety guardrails, versioning, and production readiness.
- Recruit, coach, and retain engineers while managing performance and developing senior technical leaders.
- Own production outcomes, including reliability, security, cost efficiency, supportability, and delivery predictability.
Required Qualifications
- Bachelor's, Master's, or Ph.D. in Computer Science, AI/ML, Engineering, or a related field, or equivalent experience.
- 8+ years of software engineering experience, including ownership of production systems.
- 2+ years of engineering management experience, including hiring, coaching, performance management, and delivery ownership.
- Proven ability to lead teams while remaining technically engaged in design, coding, reviews, debugging, and operations.
- Deep experience with distributed systems, cloud platforms, or AI/ML infrastructure.
- Hands-on experience building AI agents, autonomous workflows, tool-using systems, or multi-step orchestration.
- Experience with frameworks such as LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar tools.
- Strong understanding of LLM patterns, including tool calling, RAG, memory, context management, evaluation, and safety.
- Strong Python skills and experience with Kubernetes, Docker, observability, scalability, and fault tolerance.
- Strong understanding of AI security, governance, access control, auditability, and operational risk.
- Excellent communication and cross-functional leadership skills.
Preferred Qualifications
- Experience managing teams that build AI platforms, agent runtimes, inference systems, or developer platforms.
- Experience with GPU inference optimization, model serving, workflow engines, or multi-tenant cloud services.
- Experience integrating AI systems with enterprise APIs, databases, identity systems, vector stores, and policy layers.
- Experience with agent evaluation, adversarial testing, regression gates, and production observability.
- Experience using AI-assisted development tools such as Codex, Claude Code, Cursor, or Copilot.
- Experience in enterprise, cloud infrastructure, regulated, or mission-critical environments.
- Experience in data science and applied machine learning, including classical ML techniques, deep learning models, model evaluation, and production deployment.
QualificationsDisclaimer:
Range and benefit information provided in this posting are specific to the stated locations onlyUS: Hiring Range in USD from: $146,300 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - M3