Oracle Corporation

Lead Principal Incident Manager - Data Centers (on-site Nashville, TN)

Oracle Corporation • $110K — $234K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6-10+ years in incident management or infrastructure reliability engineering.
  • Experience leading high-severity incidents in large-scale environments.
  • Strong understanding of data center operations, including power and networking.
  • Proven ability to lead cross-functional teams under pressure.
  • Familiarity with incident-management frameworks like ITIL or SRE practices.
  • Bachelor's degree in a technical field preferred, or equivalent experience.

Responsibilities

  • Lead responses to critical incidents affecting OCI data center infrastructure.
  • Create and maintain incident timelines, decision logs, and action plans.
  • Provide updates to technical teams and executive stakeholders during incidents.
  • Collaborate with various teams to align responses across infrastructure layers.
  • Conduct post-incident reviews and root cause analyses for major events.
  • Analyze incident data to identify reliability risks and improvement opportunities.
  • Enhance incident-management processes and develop operational runbooks.

Benefits

  • Medical, dental, and vision insurance.
  • Short and long-term disability coverage.
  • Life insurance and AD&D.
  • Flexible Spending Accounts for healthcare and dependent care.
  • 401(k) plan with company match.
  • Flexible vacation and paid time off policies.
  • Paid parental leave and adoption assistance.
Full Job Description
Job Description

>>This position will be full-time on-site at Oracle's offices located in Nashville, TN. Relocation assistance may be available in accordance with Oracle's relocation policies. Candidates should expect a minimum of 25% travel, with additional travel as business needs require.<<

As a Lead Principal Incident Manager, you will lead the response to critical service-impacting events across Oracle Cloud Infrastructure (OCI) data center infrastructure. You will serve as a senior individual contributor and trusted operational leader, coordinating Site Operations, infrastructure engineering, network engineering, platform reliability, security, and vendor teams to restore service quickly, communicate clearly, and strengthen long-term operational resilience.

Responsibilities

Key Responsibilities

  • Lead the response to critical incidents affecting OCI data center infrastructure, including compute, storage, network, platform, and facility-dependent services.


  • Create and maintain clear incident timelines, decision logs, action plans, and stakeholder communications throughout the incident lifecycle.


  • Provide concise, accurate updates to technical teams, operational leadership, and executive stakeholders during active incidents and recovery efforts.


  • Partner with Site Operations, Infrastructure Engineering, Network Engineering, Reliability Engineering, Security Operations, and external vendors to align responses across multiple infrastructure layers.


  • Lead post-incident reviews and root cause analysis for major or recurring events, ensuring corrective actions are specified, owned, and tracked to completion.


  • Analyze incident data, operational trends, and recurring failure patterns to identify systemic reliability risks and prioritize improvement opportunities.


  • Improve incident-management processes, escalation paths, response playbooks, and operating rhythms in collaboration with technical support and engineering teams.


  • Develop and maintain operational runbooks, incident command standards, and communication templates that support disciplined, repeatable response.


  • Support implementation of automation, monitoring, observability, and health-reporting improvements that reduce detection and restoration time.


  • Track and report key operational metrics, including time to detect, time to mitigate, time to resolve, and incident recurrence.


Ideal Candidate Profile

  • 6-10+ years of experience in incident management, infrastructure reliability engineering, data center operations, or other uptime-critical environments.


  • Experience leading high-severity incidents in large-scale distributed infrastructure, cloud, data center, or hybrid environments.


  • Strong understanding of data center operations, power, cooling, networking, and storage infrastructure.


  • Demonstrated ability to lead cross-functional teams through high-pressure operational events while maintaining sound judgment and clear accountability.


  • Experience with incident-management frameworks such as ITIL, SRE practices, the Incident Command System, or equivalent operational models.


  • Bachelor's degree in a technical field preferred; equivalent experience also valued.


Skills and Competencies

  • Strong incident command, operational decision-making, and cross-team coordination capability.


  • Clear written and verbal communication skills, including the ability to deliver concise executive-level incident summaries.


  • Strong analytical approach to root cause analysis, corrective action development, and trend identification.


  • Ability to influence technical and operational decisions without formal authority.


  • Ability to translate incident learnings into durable improvements in reliability, response readiness, and operational maturity.


Preferred Skills / Certifications

  • Experience supporting AI, HPC, or large-scale GPU infrastructure, including high-density compute environments and high-performance networks.


  • Experience with CAPA reviews in JIRA environment is strongly preferred, within the data center vertical.


  • Background in facility site reliability engineering, infrastructure engineering, data center operations, or mission-critical facilities operations.


  • Experience with incident tracking, alerting, and observability tools such as PagerDuty, ServiceNow, Jira, Datadog, Prometheus, or Grafana.


  • Experience in hyperscale or colocation data center environments and with external hardware or facility vendors.


  • Formal training or certification in incident command, ITIL, service management, or related operational disciplines is a plus.


Physical Demands / Work Environment

This role supports mission-critical cloud infrastructure and may require participation in on-call incident response, site engagement, operational reviews, and travel as business needs require. You must be able to work safely in active data centers and industrial environments, with or without reasonable accommodation.

#LI-SB36

Qualifications

US: Hiring Range in USD from: $110,200 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.

About Oracle Corporation

Oracle Dyn Global Business Unit is a pioneer in managed DNS and a leader in cloud-based infrastructure that connects users with digital content and experiences across a global internet. Dyn's solution is powered by a global network that drives 40 billion traffic optimization decisions daily for more than 3,500 enterprise customers, including preeminent digital brands such as Netflix, Twitter, Linkedin and CNBC. Adding Dyn's best-in-class DNS and email services extend the Oracle cloud computing platform and provides enterprise customers with a one-stop shop for Infrastructure-as-a-Service (IaaS) and Platform-as-a-Service (PaaS). On January 31, 2017 Oracle completed the acquisition of Dyn, which now operates as an Oracle Infrastructure-as-a-Service (IaaS) global business unit (GBU).

Oracle Corporation Careers

Join Oracle Corporation, a global leader in technology and innovation, and be part of a team that values professional growth, leadership, and diversity. At Oracle, we offer unparalleled job opportunities in the tech industry, fostering a culture of innovation and continuous improvement.

Work You’ll Do

At Oracle, your work will directly impact the future of technology across industries. As part of our team, you will lead projects that redefine the way businesses operate, leveraging Oracle’s cutting-edge technology solutions. Our commitment to leadership in the tech community means you’ll be working at the forefront of innovation, enhancing your skills through hands-on experience and comprehensive diversity training.

Join Our Dynamic Team

Oracle is not just a technology company; we are a team of dedicated professionals committed to creating a supportive and inclusive environment. Here, every team member’s contribution is valued, and diversity is celebrated. With Oracle, you are not just accepting a job; you are joining a community that promotes personal and professional growth through constant learning and development opportunities.

Innovative Work and Career Advancement

Embrace the chance to do innovative work with Oracle Corporation, where we push the boundaries of what is possible. With over 130,000 dedicated professionals globally, Oracle offers a workplace where innovation and thought leadership thrive. This environment is perfect for those who are driven to explore new ideas and are eager for opportunities to advance their careers.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional looking for your next career challenge or a student seeking a promising internship, Oracle provides a range of opportunities. Explore positions that match your skills and interests in areas such as cloud computing, enterprise software, and business analytics. Our hiring process is designed to find not just the right skills but also the right fit for Oracle’s unique culture.

Benefits and Culture

Oracle is committed to supporting our employees’ life and work ambitions. We offer competitive benefits, including health insurance, retirement plans, and wellness programs, all designed to support your career and well-being. Our culture of empowerment encourages networking and collaboration across teams and geographies, ensuring that innovation and creativity flourish.

Develop Your Skills Through Training and Networking

Prepare for your future with Oracle’s comprehensive training programs. From leadership development to technical skills enhancement, we provide the tools necessary to succeed in your career and stay ahead in the industry. Networking within Oracle’s global community will also open doors to collaborative opportunities and career advancement.

Stay Connected with Oracle Careers

Keep up to date with the latest from Oracle Corporation by following our careers blog. Gain insights from the experts and learn about new job openings as they become available. Personalize your job search and stay informed about Oracle’s career events and professional development opportunities.

Join Oracle Corporation—Where Careers Grow

At Oracle, we believe in nurturing the potential of our employees. The growth of our company is driven by the individual successes of our team members. We invite you to bring your unique talents to Oracle, join our mission to drive technological innovation, and help shape the future of the digital world.

Search Oracle Jobs

Ready to take the next step in your career? Search for open positions that align with your skills and passions. We are continuously looking for curious, creative, and motivated individuals to join our team. Explore the opportunities and find out how you can contribute to the success of Oracle Corporation.

Oracle Corporation: Leadership, Innovation, Opportunity.

Learn more about Oracle Corporation
Size
143,000 employees
Market Cap
$217.3 billion
Industry
Net Income
$12.8 billion
Founded
1977
5 Year Trend
+2.3%
Revenue
$39.6 billion
NASDAQ

Similar Jobs

More Jobs at Oracle Corporation

More Information Technology Jobs

Find similar Lead Principal Incident Manager - Data Centers (on-site Nashville, TN) jobs: