Oracle Corporation

Director, Reliability Engineering (Nashville, TN on-site)

Oracle Corporation • $146K — $306K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years in engineering or infrastructure, notably in mission-critical environments.
  • 5+ years of leadership experience managing technical teams across multiple locations.
  • Expertise in electrical distribution, UPS systems, generators, and cooling systems.
  • Strong foundation in reliability engineering methodologies like FMEA, RCA, and RCM.
  • Proven ability to use data to influence engineering decisions and identify systemic risks.

Responsibilities

  • Lead the Reliability Engineering team across OCI's data center portfolio.
  • Define a multi-year reliability strategy that enhances infrastructure resilience.
  • Establish reliability engineering standards and methodologies organization-wide.
  • Develop a high-performing team with clear objectives and accountability.
  • Create programs to boost reliability for critical infrastructure assets.
  • Implement a reliability measurement framework for asset health and performance tracking.
  • Partner with various departments to incorporate operational insights into long-term strategies.

Benefits

  • Potential for relocation assistance in line with company policies.
  • Opportunity for professional development and career advancement.
  • Exposure to advanced engineering methodologies and cutting-edge technologies.
  • Participation in a global, mission-critical infrastructure portfolio.
Full Job Description
Job Description

>>This position will be full-time on-site at Oracle's offices located in Nashville, TN. Relocation assistance may be available in accordance with Oracle's relocation policies. Candidates should expect a minimum of 25% travel, with additional travel as business needs require.<<

As Director of Building Automation, you will provide strategic and organizational leadership for the reliability engineering function supporting Oracle Cloud Infrastructure's mission-critical data center portfolio. You will own the vision, operating model, engineering standards, and portfolio programs that improve infrastructure availability, maintainability, resilience, and lifecycle performance at scale.

This role requires significant hands-on and leadership experience within mission-critical environments. Direct data center experience is preferred. Candidates must have demonstrated experience supporting the electrical, mechanical, controls, and operational systems required to maintain continuous operations in mission-critical environments.

You will lead managers, engineers, analysts, and technical programs responsible for reliability engineering, asset performance, predictive maintenance, failure analysis, defect elimination, and lifecycle risk management across OCI's data center infrastructure.

You will partner with senior leaders across Data Center Operations, Engineering, Design, Construction, Commissioning, Automation, Procurement, and other infrastructure organizations to translate operational experience and engineering data into long-term reliability strategy. You will ensure lessons learned from incidents, equipment performance, maintenance activities, and portfolio trends result in durable improvements to standards, designs, operating practices, and investment priorities.

Success in this role requires the ability to operate at both strategic and technical levels-setting multi-year direction for the reliability organization while maintaining sufficient engineering depth and data center operational knowledge to challenge assumptions, assess complex infrastructure risks, and drive disciplined decision-making across a rapidly scaling global portfolio.

Responsibilities

Key Responsibilities
  • Lead the Reliability Engineering organization supporting multiple regions, sites, and infrastructure programs across OCI's mission-critical data center portfolio.
  • Define the multi-year reliability engineering strategy, organizational roadmap, operating model, and investment priorities required to improve data center infrastructure resilience and support OCI's continued growth.
  • Establish and govern portfolio-wide reliability engineering standards and methodologies, including FMEA/FMECA, RCA, Reliability-Centered Maintenance (RCM), criticality assessment, defect elimination, reliability growth, and continuous improvement practices.
  • Build and lead a high-performing organization of managers, engineers, analysts, and technical specialists, establishing clear accountability, technical expectations, career development, and succession plans.
  • Own portfolio-level programs that improve reliability across critical data center infrastructure, including electrical distribution, UPS systems, generators, mechanical cooling systems, controls, automation, and supporting facility systems.
  • Establish a comprehensive reliability measurement framework that provides leadership with visibility into asset health, failure trends, systemic risks, repeat events, corrective actions, maintenance effectiveness, and lifecycle exposure.
  • Define reliability KPIs, targets, governance mechanisms, and executive reporting that enable data-driven prioritization of operational and engineering investments.
  • Establish governance for corrective and preventive actions resulting from incidents, root cause analyses, audits, equipment failures, and reliability trend reviews, ensuring actions are completed, verified for effectiveness, and sustained.
  • Drive systematic identification and elimination of recurring and systemic failure modes across the data center portfolio rather than relying solely on site-specific remediation.
  • Sponsor the development and adoption of predictive and condition-based maintenance capabilities, including monitoring, analytics, automation, asset health modeling, and emerging technologies that improve early detection of equipment degradation and failure risk.
  • Partner with Data Center Operations leadership to continuously improve maintenance strategy, operational readiness, troubleshooting practices, procedures, failure response, and infrastructure risk management.
  • Provide reliability governance and technical leadership for commissioning, acceptance testing, operational handover, major maintenance, retrofits, capacity expansion, and infrastructure lifecycle decisions.
  • Establish portfolio approaches to asset lifecycle management, including equipment health, utilization, failure history, remaining useful life, obsolescence, spare parts strategy, replacement planning, and end-of-life risk.
  • Partner with Design, Construction, Engineering, and Procurement leadership to ensure lessons from operating facilities influence equipment specifications, design standards, redundancy strategies, maintainability requirements, vendor selection, and total cost of ownership.
  • Develop mechanisms to convert site-level events and engineering findings into portfolio-wide standards, design changes, maintenance improvements, and risk-reduction programs.
  • Lead technical and business reviews of significant reliability risks and provide clear recommendations regarding mitigation strategies, priorities, investment requirements, and residual operational risk.
  • Develop strong partnerships with equipment manufacturers, service providers, and technology partners to improve equipment performance, failure intelligence, serviceability, and long-term reliability.
  • Establish effective operating rhythms for the organization, including portfolio reviews, technical reviews, risk escalation, program governance, resource prioritization, and executive communications.
  • Represent Reliability Engineering in senior leadership discussions involving infrastructure risk, operational performance, capacity growth, capital planning, and long-term data center strategy.

Minimum Qualifications
  • 10+ years of progressive engineering, reliability, maintenance, critical facilities, or infrastructure experience, including significant experience directly supporting mission-critical environments.
  • Demonstrated experience working within mission-critical operations or engineering environments where infrastructure availability, redundancy, maintenance execution, and operational risk directly affect service continuity.
  • 5+ years of progressive leadership experience, including responsibility for engineering managers, senior technical professionals, or large multi-site technical organizations and programs.
  • Demonstrated technical knowledge of mission-critical infrastructure, including experience with electrical distribution, UPS systems, generators, mechanical cooling systems, controls/automation, and integrated facility operations.
  • Demonstrated experience developing and implementing reliability, maintenance, asset-management, or operational excellence strategies across mission-critical infrastructure.
  • Strong working knowledge of reliability engineering methodologies, including structured root cause analysis, FMEA/FMECA, RCM, criticality analysis, defect elimination, reliability metrics, and lifecycle risk management.
  • Demonstrated experience evaluating infrastructure failures, operational events, equipment performance, maintenance effectiveness, and systemic reliability risks within mission-critical environments.
  • Demonstrated ability to use operational and engineering data to identify systemic risks, establish priorities, and influence significant technical or business decisions.
  • Experience leading cross-functional initiatives involving Data Center Operations, Engineering, Design, Construction, Commissioning, Procurement, OEMs, vendors, and other technical stakeholders.
  • Experience establishing engineering governance, standards, KPIs, and management mechanisms across multiple sites, regions, or infrastructure programs in mission-critical environments.
  • Demonstrated ability to communicate complex technical risks, tradeoffs, and investment recommendations to senior and executive leadership.
  • Bachelor's degree in Electrical Engineering, Mechanical Engineering, Industrial Engineering, Systems Engineering, Reliability Engineering, or a related technical discipline; or equivalent relevant industry experience.

Skills and Competencies
  • Mission-Critical Technical Leadership: Strong understanding of mission-critical infrastructure, operating practices, redundancy, maintenance risk, failure modes, and the interdependencies between electrical, mechanical, controls, and operational systems.
  • Organizational Leadership: Ability to build, develop, and lead managers and senior technical professionals while establishing clear accountability and a strong engineering culture.
  • Reliability Strategy: Ability to translate data center infrastructure performance, business growth, and operational risk into a coherent multi-year reliability strategy and investment roadmap.
  • Technical Judgment: Ability to evaluate complex infrastructure reliability issues, challenge technical assumptions, understand operational consequences, and make decisions under uncertainty.
  • Systems Thinking: Ability to connect individual equipment failures and site-level events to systemic portfolio risks involving design, maintenance, operations, suppliers, processes, or organizational practices.
  • Data-Driven Decision Making: Ability to convert large volumes of operational, maintenance, failure, and asset data into actionable insights and investment priorities.
  • Executive Influence: Ability to communicate technical risk and recommendations clearly to senior leaders and build alignment across organizations with different objectives and priorities.
  • Operational Excellence: Strong commitment to disciplined execution, corrective-action rigor, standardization, measurable improvement, and sustained results.
  • Change Leadership: Ability to introduce and scale new engineering methods, technologies, processes, and operating models across a large and geographically distributed data center organization.
  • Talent Development: Demonstrated ability to develop engineering leaders and technical talent, establish career paths, strengthen organizational capability, and build succession depth.

Preferred Qualifications
  • Direct data center experience supporting mission-critical infrastructure and operations is preferred.
  • Experience leading reliability engineering, critical facilities engineering, or asset-management organizations within hyperscale, colocation, or large-scale enterprise data centers.
  • Experience supporting geographically distributed or global data center portfolios.
  • Deep technical expertise in one or more critical infrastructure domains, with broad working knowledge across electrical distribution, UPS, standby generation, mechanical cooling, controls/automation, and integrated facility operations.
  • Experience developing and scaling predictive maintenance, condition-based monitoring, failure trend analysis, asset health modeling, and equipment risk-ranking programs within data center environments.
  • Advanced knowledge of reliability, availability, and maintainability analysis; maintenance strategy optimization; spare parts planning; lifecycle modeling; and total cost of ownership.
  • Experience governing commissioning, operational acceptance, maintenance program design, and readiness of new or modified mission-critical data center infrastructure.
  • Experience with CMMS/EAM, DCIM, EPMS, BMS, monitoring, telemetry, analytics, and automation platforms used to manage critical data center infrastructure.
  • Experience developing portfolio-level KPI frameworks, executive dashboards, reliability reviews, and risk-governance mechanisms.
  • Experience partnering with OEMs and strategic suppliers to address systemic equipment issues, improve product reliability, and influence equipment roadmaps or specifications.
  • Experience incorporating operational lessons learned into engineering standards, design requirements, equipment specifications, and capital investment decisions.
  • Demonstrated experience managing organizational growth, workforce planning, resource prioritization, and technical capability development across geographically distributed teams.

Preferred Credentials / Certifications
  • Certified Maintenance & Reliability Professional (CMRP) preferred.
  • Certified Reliability Engineer (CRE) preferred.
  • ASQ, SMRP, or equivalent reliability, maintenance, engineering, or quality certifications are a plus.
  • Data center or critical-environment credentials, including relevant Uptime Institute training or certifications, are a plus.
  • OEM, controls, analytics, asset-management, or condition-monitoring training applicable to critical data center infrastructure is a plus.
  • Advanced training or certification in FMEA/FMECA, RCA, RCM, Lean, Six Sigma, or structured problem-solving methodologies is a plus.
  • Advanced technical or business degree is a plus.

Leadership Scope
  • The Director - Reliability Engineering is expected to operate beyond individual programs or regions and establish the mechanisms through which OCI manages reliability risk at portfolio scale. The role requires balancing immediate operational priorities with long-term infrastructure strategy and ensuring reliability engineering becomes an increasingly predictive, data-driven, and standardized capability.
  • The Director will be accountable for building an organization capable of identifying emerging reliability risks before they become widespread operational issues, converting lessons learned into durable improvements, and e

About Oracle Corporation

Oracle Dyn Global Business Unit is a pioneer in managed DNS and a leader in cloud-based infrastructure that connects users with digital content and experiences across a global internet. Dyn's solution is powered by a global network that drives 40 billion traffic optimization decisions daily for more than 3,500 enterprise customers, including preeminent digital brands such as Netflix, Twitter, Linkedin and CNBC. Adding Dyn's best-in-class DNS and email services extend the Oracle cloud computing platform and provides enterprise customers with a one-stop shop for Infrastructure-as-a-Service (IaaS) and Platform-as-a-Service (PaaS). On January 31, 2017 Oracle completed the acquisition of Dyn, which now operates as an Oracle Infrastructure-as-a-Service (IaaS) global business unit (GBU).

Oracle Corporation Careers

Join Oracle Corporation, a global leader in technology and innovation, and be part of a team that values professional growth, leadership, and diversity. At Oracle, we offer unparalleled job opportunities in the tech industry, fostering a culture of innovation and continuous improvement.

Work You’ll Do

At Oracle, your work will directly impact the future of technology across industries. As part of our team, you will lead projects that redefine the way businesses operate, leveraging Oracle’s cutting-edge technology solutions. Our commitment to leadership in the tech community means you’ll be working at the forefront of innovation, enhancing your skills through hands-on experience and comprehensive diversity training.

Join Our Dynamic Team

Oracle is not just a technology company; we are a team of dedicated professionals committed to creating a supportive and inclusive environment. Here, every team member’s contribution is valued, and diversity is celebrated. With Oracle, you are not just accepting a job; you are joining a community that promotes personal and professional growth through constant learning and development opportunities.

Innovative Work and Career Advancement

Embrace the chance to do innovative work with Oracle Corporation, where we push the boundaries of what is possible. With over 130,000 dedicated professionals globally, Oracle offers a workplace where innovation and thought leadership thrive. This environment is perfect for those who are driven to explore new ideas and are eager for opportunities to advance their careers.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional looking for your next career challenge or a student seeking a promising internship, Oracle provides a range of opportunities. Explore positions that match your skills and interests in areas such as cloud computing, enterprise software, and business analytics. Our hiring process is designed to find not just the right skills but also the right fit for Oracle’s unique culture.

Benefits and Culture

Oracle is committed to supporting our employees’ life and work ambitions. We offer competitive benefits, including health insurance, retirement plans, and wellness programs, all designed to support your career and well-being. Our culture of empowerment encourages networking and collaboration across teams and geographies, ensuring that innovation and creativity flourish.

Develop Your Skills Through Training and Networking

Prepare for your future with Oracle’s comprehensive training programs. From leadership development to technical skills enhancement, we provide the tools necessary to succeed in your career and stay ahead in the industry. Networking within Oracle’s global community will also open doors to collaborative opportunities and career advancement.

Stay Connected with Oracle Careers

Keep up to date with the latest from Oracle Corporation by following our careers blog. Gain insights from the experts and learn about new job openings as they become available. Personalize your job search and stay informed about Oracle’s career events and professional development opportunities.

Join Oracle Corporation—Where Careers Grow

At Oracle, we believe in nurturing the potential of our employees. The growth of our company is driven by the individual successes of our team members. We invite you to bring your unique talents to Oracle, join our mission to drive technological innovation, and help shape the future of the digital world.

Search Oracle Jobs

Ready to take the next step in your career? Search for open positions that align with your skills and passions. We are continuously looking for curious, creative, and motivated individuals to join our team. Explore the opportunities and find out how you can contribute to the success of Oracle Corporation.

Oracle Corporation: Leadership, Innovation, Opportunity.

Learn more about Oracle Corporation
Size
143,000 employees
Market Cap
$217.3 billion
Industry
Net Income
$12.8 billion
Founded
1977
5 Year Trend
+2.3%
Revenue
$39.6 billion
NASDAQ

Similar Jobs

More Jobs at Oracle Corporation

More Enterprise Technology Jobs

Find similar Director, Reliability Engineering (Nashville, TN on-site) jobs: