Job DescriptionJob Description
Oracle Cloud Infrastructure (OCI) is building some of the world's largest and most advanced GPU clusters to power the next generation of AI. The Strategic Customer Engineering (SCE) Core Infrastructure team-also known as the AI/ML Forward Deployed Infrastructure Engineering team-provides white-glove engineering and operational support to OCI's most strategic AI/ML infrastructure customers.
As a trusted partner to our customers, we play a critical role in designing, deploying, operating, and optimizing the infrastructure that powers some of the largest and most demanding GPU and AI/ML environments in the world. Our team works closely with customers and internal engineering organizations to ensure exceptional reliability, performance, and scalability for mission-critical AI workloads.
We are seeking a Manager to lead and develop a class of early-career Infrastructure Engineers. This leader will create the conditions for new professionals to build sound technical judgment, strong customer habits, and the practical skills required to support strategic AI/ML infrastructure customers on OCI. The manager will combine people leadership with hands-on operational guidance, helping the class grow into effective engineers who can independently contribute to customer execution, automation, troubleshooting, and infrastructure performance.
#LI-ES2
ResponsibilitiesLead, mentor, and develop a class of early-career Engineers, setting clear expectations for technical growth, customer engagement, quality, and ownership.
Design and run an onboarding and development program that gives early-career professionals progressive experience with GPU cluster deployment, Slurm and/or Oracle Kubernetes Engine (OKE), automation, monitoring, multicloud, database and operational readiness.
Provide regular coaching, feedback, career guidance, and practical learning opportunities. Pair team members with appropriate mentors and create stretch assignments that build confidence and sound engineering judgment.
Guide the team in building collaborative relationships with OCI Services, customers, sales, and core engineering teams. Model clear technical communication and help team members become effective, trusted customer partners.
Establish a safe, accountable environment where early-career professionals learn to take ownership of problems, investigate issues thoroughly, document risks, and escalate with context when needed.
Coach team members on infrastructure performance fundamentals, including parameter tuning, resource utilization, caching, and data pre-processing techniques for AI/ML workloads.
Oversee the team's approach to troubleshooting performance, scalability, and reliability issues. Review complex cases, remove barriers, and ensure lessons learned are incorporated into team practices.
Ensure the team documents infrastructure designs, configurations, runbooks, and procedures so knowledge is shared, work is maintainable, and new professionals can learn from repeatable practices.
Help team members develop into credible customer advocates who can explain advanced GPU solution best practices and support customers and partners as they migrate workloads to the cloud.
Partner with senior technical leaders to calibrate the class's development progress, staffing needs, and readiness for increasingly independent customer-facing and engineering responsibilities.
Qualifications:7+ years of relevant infrastructure, cloud, platform, or software engineering experience, including experience leading or mentoring engineers.
Strong communication and collaboration skills, with the ability to coach early-career professionals and translate technical concepts for customers, partners, and non-technical stakeholders.
Demonstrated people-management, mentoring, or talent-development skills, with a thoughtful and consistent approach to feedback and performance management.
Experience building, operating, or supporting distributed and cloud infrastructure solutions in a customer-focused environment.
Demonstrated ability to:
- Develop short-, medium-, and long-term development plans that connect individual growth to team and business objectives.
- Work across functional areas with senior leaders and technical partners to secure opportunities, support, and clear expectations for an early-career class.
- Influence through coaching, sound judgment, and clear communication in situations involving competing priorities or sensitive feedback.
BS or MS degree, or equivalent experience, relevant to the functional area.
Working knowledge of infrastructure automation and orchestration tools such as Ansible, Terraform, Python, Docker, Kubernetes, and related tooling.
Solid understanding of networking concepts, security principles, and operational best practices.
Excellent problem-solving skills, with the ability to guide others through structured troubleshooting and drive timely resolution in a fast-paced environment.
Strong Linux skills with hands-on experience in Oracle Linux/RHEL/CentOS, Ubuntu, or Debian, including system administration, package management, shell scripting, and performance optimization.
Proficiency in at least one programming language, such as Python, Rust, Go, Java, or Scala, and the ability to use code reviews and pairing to develop others.
Experience designing, implementing, managing, or supporting infrastructure for AI/ML, HPC, GPU, or similarly complex cloud workloads.
Core ResponsibilitiesPlanning & Execution:Plan and guide the class's work across onboarding, customer support, automation, and development goals. Set clear priorities, assign work that matches each person's readiness, monitor delivery, and adjust learning plans as business needs change.
Collaboration & Partnership:Build productive cross-functional partnerships that give early-career professionals access to the expertise and context they need. Model inclusive collaboration, actively seek diverse perspectives, and help the team communicate expectations clearly with stakeholders and customers.
Problem Solving:Coach the team to analyze operational and technical issues using sound problem-solving practices. Review complex or ambiguous cases, help identify root causes, and turn lessons learned into guidance that prevents recurrence.
Continuous Learning:Create structured opportunities for team members to build expertise through training, paired work, technical reviews, and hands-on customer scenarios. Track individual skill growth, identify gaps early, and reinforce a culture of knowledge sharing and continuous learning.
Continuous Improvement:Enable the team to improve processes, protocols, runbooks, and workflows. Gather feedback from early-career professionals and partners, prioritize practical improvements, and help the team learn how to turn ideas into durable operational practices.
Performance and Development:Drive consistent performance and development through timely, specific feedback, regular coaching, and clear growth expectations. Support hiring, onboarding, talent reviews, and promotion readiness in partnership with leadership and HR, while aligning individual development plans to organizational needs.
Qualifications