Oracle Corporation

Principal Systems Engineer

Oracle Corporation • $130K — $155K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in software operations or infrastructure automation, proficient in Python and Bash.
  • Expertise in Linux administration (Ubuntu, Oracle Linux) in production environments.
  • Strong knowledge of distributed systems communication patterns.
  • Experience with data-center and host-lifecycle operations, including provisioning and fleet recovery.
  • Excellent problem-solving and troubleshooting skills.
  • Strong communication and teamwork abilities.
  • Familiarity with observability tools for monitoring and alerting.
  • Experience with AI agents and tooling.

Responsibilities

  • Build and maintain automation and operational tooling for GPU infrastructure.
  • Collaborate with engineering and operations teams to ensure GPU fleet availability.
  • Enhance monitoring, alerting, and diagnostics for fleet performance and health.
  • Act as the senior escalation point for complex GPU issues.
  • Conduct incident responses and root-cause analyses.
  • Improve AI2 Ops processes and GPU fleet readiness.
  • Document operational procedures and create runbooks.

Benefits

  • Opportunity to work on cutting-edge GPU technology in a cloud environment.
  • Collaborative work environment with cross-team engagement.
  • Access to continuous learning and mentorship opportunities.
  • Participation in on-call rotations for skill enhancement.
  • Impactful contributions to the AI Infra Operations team.
Full Job Description
Job Description

We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Systems Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational tooling for GPU fleets across multiple regions, ensuring high availability. You will collaborate closely with engineering and operations teams to build robust automation, observability, and reliability solutions while continuously improving GPU operations.

Responsibilities

  • Build and maintain automation and operational tooling for OCI GPU infrastructure across multiple geographic regions.
  • Drive collaboration with software engineers, hardware teams, and operations partners to maintain a highly available GPU fleet.
  • Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using tools such as Grafana.
  • Serve as the senior escalation point for complex GPU host and repair issues
  • Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.
  • Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.
  • Participate in on-call rotations and provide support for critical infrastructure issues.
  • Document operational procedures, automation workflows, troubleshooting guides, and runbooks.
  • Build and improve AI agents, ensuring safe rollout, execution and monitoring.
  • Mentor and guide junior engineers in operational best practices, provide senior technical support, and drive constant improvement.


Required Qualifications:

  • 8+ years of software operations or infrastructure automation experience with strong proficiency in Python and Bash.
  • Expert Linux administration experience, particularly Ubuntu and Oracle Linux, in large-scale production environments.
  • Strong understanding of distributed systems, including peer-to-peer, node-to-node, and service-to-service communication patterns.
  • Strong data-center and host-lifecycle experience, including provisioning, validation, repair workflows, hardware replacement, and fleet recovery.
  • Strong problem-solving and troubleshooting skills.
  • Excellent communication and teamwork skills.
  • Experience with observability tooling, including metrics, logging, dashboards, and alerting.
  • Experience with AI agents and tooling
  • Experience leading on-call operations and incident response.
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.


Preferred Skills:

  • Hands-on experience with GPU infrastructure, including NVIDIA and AMD based systems, and familiarity with GPU shapes or instance families.
  • Experience operating or automating GPU, compute, or other large-scale cloud infrastructure.


Qualifications

Minimum Job Qualifications
Education and/or Experience:
11 years of experience in computer network administration, database management, systems architecture, systems administration, or related field

OR

Bachelor's Degree in information technology, computer science, engineering or related field AND 7 years of experience in computer network administration, database management, systems architecture, systems administration, or related field

OR

Master's Degree in information technology, computer science, engineering or related field AND 5 years of experience in computer network administration, database management, systems architecture, systems administration, or related field.

Job Skills:
Same skills as prior level plus;
Cybersecurity Trends Demonstrated ability in or knowledge of cybersecurity trends, including staying current with industry threats, best practices, and emerging technologies.
Incident Management and Response Demonstrated ability in or knowledge of incident management and response, including timely handling and escalation of incidents to minimize business impact.
Training and Development Demonstrated ability to design and deliver effective training programs to build team and individual capabilities.
Technical Account Management Demonstrated ability to manage technical relationships and ensure successful outcomes with key accounts.
Cloud Architecture Demonstrated ability in or knowledge of cloud architecture, including designing scalable, reliable, and performant cloud services.

Preferred Job Qualifications
Education and/or Experience:
12 years of experience in computer network administration, database management, systems architecture, systems administration, or related field

OR

Bachelor's Degree in information technology, computer science, engineering or related field AND 8 years of experience in computer network administration, database management, systems architecture, systems administration, or related field

OR

Master's Degree in information technology, computer science, engineering or related field AND 6 years of experience in computer network administration, database management, systems architecture, systems administration, or related field.

Job Skills:
Same skills as prior level

About Oracle Corporation

Oracle Dyn Global Business Unit is a pioneer in managed DNS and a leader in cloud-based infrastructure that connects users with digital content and experiences across a global internet. Dyn's solution is powered by a global network that drives 40 billion traffic optimization decisions daily for more than 3,500 enterprise customers, including preeminent digital brands such as Netflix, Twitter, Linkedin and CNBC. Adding Dyn's best-in-class DNS and email services extend the Oracle cloud computing platform and provides enterprise customers with a one-stop shop for Infrastructure-as-a-Service (IaaS) and Platform-as-a-Service (PaaS). On January 31, 2017 Oracle completed the acquisition of Dyn, which now operates as an Oracle Infrastructure-as-a-Service (IaaS) global business unit (GBU).

Oracle Corporation Careers

Join Oracle Corporation, a global leader in technology and innovation, and be part of a team that values professional growth, leadership, and diversity. At Oracle, we offer unparalleled job opportunities in the tech industry, fostering a culture of innovation and continuous improvement.

Work You’ll Do

At Oracle, your work will directly impact the future of technology across industries. As part of our team, you will lead projects that redefine the way businesses operate, leveraging Oracle’s cutting-edge technology solutions. Our commitment to leadership in the tech community means you’ll be working at the forefront of innovation, enhancing your skills through hands-on experience and comprehensive diversity training.

Join Our Dynamic Team

Oracle is not just a technology company; we are a team of dedicated professionals committed to creating a supportive and inclusive environment. Here, every team member’s contribution is valued, and diversity is celebrated. With Oracle, you are not just accepting a job; you are joining a community that promotes personal and professional growth through constant learning and development opportunities.

Innovative Work and Career Advancement

Embrace the chance to do innovative work with Oracle Corporation, where we push the boundaries of what is possible. With over 130,000 dedicated professionals globally, Oracle offers a workplace where innovation and thought leadership thrive. This environment is perfect for those who are driven to explore new ideas and are eager for opportunities to advance their careers.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional looking for your next career challenge or a student seeking a promising internship, Oracle provides a range of opportunities. Explore positions that match your skills and interests in areas such as cloud computing, enterprise software, and business analytics. Our hiring process is designed to find not just the right skills but also the right fit for Oracle’s unique culture.

Benefits and Culture

Oracle is committed to supporting our employees’ life and work ambitions. We offer competitive benefits, including health insurance, retirement plans, and wellness programs, all designed to support your career and well-being. Our culture of empowerment encourages networking and collaboration across teams and geographies, ensuring that innovation and creativity flourish.

Develop Your Skills Through Training and Networking

Prepare for your future with Oracle’s comprehensive training programs. From leadership development to technical skills enhancement, we provide the tools necessary to succeed in your career and stay ahead in the industry. Networking within Oracle’s global community will also open doors to collaborative opportunities and career advancement.

Stay Connected with Oracle Careers

Keep up to date with the latest from Oracle Corporation by following our careers blog. Gain insights from the experts and learn about new job openings as they become available. Personalize your job search and stay informed about Oracle’s career events and professional development opportunities.

Join Oracle Corporation—Where Careers Grow

At Oracle, we believe in nurturing the potential of our employees. The growth of our company is driven by the individual successes of our team members. We invite you to bring your unique talents to Oracle, join our mission to drive technological innovation, and help shape the future of the digital world.

Search Oracle Jobs

Ready to take the next step in your career? Search for open positions that align with your skills and passions. We are continuously looking for curious, creative, and motivated individuals to join our team. Explore the opportunities and find out how you can contribute to the success of Oracle Corporation.

Oracle Corporation: Leadership, Innovation, Opportunity.

Learn more about Oracle Corporation
Size
143,000 employees
Market Cap
$217.3 billion
Industry
Net Income
$12.8 billion
Founded
1977
5 Year Trend
+2.3%
Revenue
$39.6 billion
NASDAQ

Similar Jobs

More Jobs at Oracle Corporation

More Information Technology Jobs

Find similar Principal Systems Engineer jobs: