Oracle Corporation

Lead Principal Site Reliability Engineer

Oracle Corporation$96K — $264K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in site reliability engineering or related technical fields.
  • Expertise in designing and improving high-availability production systems.
  • Proficient in cloud platforms like AWS, Azure, or GCP.
  • Strong knowledge of containerization tools like Docker and Kubernetes.
  • Experience with infrastructure-as-code tools such as Terraform or Ansible.
  • Solid programming skills in languages like Python, Go, or Java.
  • Proven leadership in managing critical production incidents.

Responsibilities

  • Define site reliability engineering strategies for large-scale platforms.
  • Establish standards for reliability and engineering practices across teams.
  • Serve as a technical authority for system reliability and performance.
  • Identify reliability risks and lead initiatives to mitigate them.
  • Mentor site reliability and software engineers in best practices.
  • Implement observability strategies to monitor system health and performance.
  • Conduct post-incident reviews to enhance operational processes.

Benefits

  • Medical, dental, and vision insurance with expert medical opinion.
  • Short and long-term disability insurance.
  • 401(k) plan with company match.
  • Flexible Vacation and paid time off policies.
  • Paid parental leave and adoption assistance.
  • Employee Stock Purchase Plan and financial planning services.
  • Voluntary benefits including pet insurance and legal assistance.
Full Job Description
Job Description

Serves as a consultant and leads the design and architecture of infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams to develop reliable and scalable infrastructures. Recommends methods for performing data collection to maintain and optimize operations and reliability. Oversees incident response and/or maintenance tasks. Provides strategic, future-oriented health and performance reporting. Contributes to strategies for automation and reviews the development and implementation of automation. Provides expert-level communication about services and anticipates, analyzes, and explains the impact of changes, considering strategic goals. Serves as a role model in providing support for technology and reviews documentation for accuracy. Leads the implementation of innovative tools and provides expertise in site reliability trends.

Responsibilities

Key Responsibilities

Reliability Strategy and Technical Leadership
  • Define and drive the site reliability engineering strategy for large-scale, distributed, and business-critical platforms.
  • Establish reliability standards, engineering practices, and operational readiness requirements across multiple teams.
  • Serve as a senior technical authority for system reliability, scalability, resilience, performance, and production operations.
  • Influence architecture and design decisions to ensure systems are supportable, observable, fault tolerant, and capable of meeting availability objectives.
  • Identify systemic reliability risks and lead cross-functional initiatives to address them.
  • Provide technical direction and mentorship to site reliability engineers, software engineers, platform engineers, and operations teams.
  • Lead technical reviews and promote consistent engineering practices across the organization.

Service Reliability and Observability
  • Define and implement service-level indicators, service-level objectives, error budgets, and operational health metrics.
  • Develop comprehensive monitoring, logging, tracing, alerting, and observability strategies.
  • Improve the quality and actionability of alerts while reducing unnecessary operational noise.
  • Establish dashboards and reporting mechanisms that clearly communicate service health, performance, capacity, and risk.
  • Use production data and reliability trends to prioritize engineering investments and continuous-improvement initiatives.

Automation and Platform Engineering
  • Design and implement automation that reduces manual intervention, operational toil, and human error.
  • Build or enhance tools for deployment, configuration management, infrastructure provisioning, incident response, and service recovery.
  • Promote infrastructure-as-code, policy-as-code, automated testing, and repeatable deployment practices.
  • Partner with development teams to improve continuous integration and continuous delivery pipelines.
  • Develop self-healing and automated remediation capabilities where appropriate.
  • Contribute production-quality software and reusable platform components using modern programming and scripting languages.

Incident Management and Problem Resolution
  • Provide technical leadership during complex, high-severity production incidents.
  • Coordinate diagnosis, containment, recovery, and stakeholder communication during service disruptions.
  • Lead blameless post-incident reviews and ensure that corrective actions address root causes rather than symptoms.
  • Identify recurring failure patterns and develop long-term engineering solutions.
  • Improve incident-management processes, escalation procedures, runbooks, and recovery playbooks.
  • Participate in an on-call rotation or provide senior escalation support for critical services, as required.

Capacity, Performance, and Resilience
  • Lead capacity planning, performance analysis, load testing, and demand forecasting for critical platforms.
  • Identify performance bottlenecks and recommend architectural or operational improvements.
  • Design and validate high-availability, disaster-recovery, backup, and business-continuity capabilities.
  • Lead resilience testing, failure-mode analysis, game days, and controlled fault-injection exercises.
  • Ensure recovery-time and recovery-point objectives are defined, tested, and achievable.

Security and Operational Governance
  • Partner with security and compliance teams to embed security into infrastructure, automation, and operational practices.
  • Support vulnerability remediation, access-control improvements, audit readiness, and secure configuration management.
  • Ensure production environments meet organizational standards for change management, data protection, and operational governance.
  • Balance reliability, security, delivery speed, cost, and business priorities when recommending technical solutions.

Required Qualifications
  • Extensive professional experience in site reliability engineering, software engineering, cloud infrastructure, platform engineering, systems engineering, or a related technical discipline.
  • Demonstrated experience designing, operating, and improving highly available production systems at significant scale.
  • Deep knowledge of distributed systems, cloud architecture, networking, operating systems, storage, databases, and service dependencies.
  • Advanced experience with at least one major cloud platform, such as Oracle Cloud Infrastructure, Amazon Web Services, Microsoft Azure, or Google Cloud Platform.
  • Strong experience with containerization and orchestration technologies, including Docker and Kubernetes.
  • Proven expertise with infrastructure-as-code and configuration-management technologies such as Terraform, Ansible, Chef, Puppet, or equivalent tools.
  • Experience implementing observability solutions using metrics, logs, traces, dashboards, and automated alerting.
  • Strong programming or scripting skills in one or more languages such as Python, Go, Java, JavaScript, Bash, or similar.
  • Experience with continuous integration, continuous delivery, automated testing, and modern release-management practices.
  • Demonstrated leadership during critical production incidents and complex technical investigations.
  • Ability to diagnose difficult system issues across applications, infrastructure, networks, databases, and cloud services.
  • Strong written and verbal communication skills, including the ability to explain technical risk and recommendations to engineering leaders and business stakeholders.
  • Proven ability to lead cross-functional technical initiatives without relying solely on formal authority.

Preferred Qualifications
  • Experience supporting enterprise-scale cloud services, software-as-a-service platforms, or other high-availability customer-facing systems.
  • Experience defining and operating service-level objectives, error budgets, and reliability scorecards.
  • Knowledge of chaos engineering, resilience testing, and automated recovery techniques.
  • Experience with multi-region, hybrid-cloud, or multi-cloud architectures.
  • Familiarity with security frameworks, compliance requirements, and regulated operating environments.
  • Experience improving cloud cost efficiency, capacity utilization, or infrastructure performance.
  • Contributions to internal engineering standards, technical communities, open-source projects, or industry publications.
  • Bachelor's or advanced degree in computer science, engineering, information systems, or a related field, or equivalent practical experience.


Qualifications

US: Hiring Range in USD from: $96,300 to $264,100 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC5

About Oracle Corporation

Oracle Dyn Global Business Unit is a pioneer in managed DNS and a leader in cloud-based infrastructure that connects users with digital content and experiences across a global internet. Dyn's solution is powered by a global network that drives 40 billion traffic optimization decisions daily for more than 3,500 enterprise customers, including preeminent digital brands such as Netflix, Twitter, Linkedin and CNBC. Adding Dyn's best-in-class DNS and email services extend the Oracle cloud computing platform and provides enterprise customers with a one-stop shop for Infrastructure-as-a-Service (IaaS) and Platform-as-a-Service (PaaS). On January 31, 2017 Oracle completed the acquisition of Dyn, which now operates as an Oracle Infrastructure-as-a-Service (IaaS) global business unit (GBU).

Oracle Corporation Careers

Join Oracle Corporation, a global leader in technology and innovation, and be part of a team that values professional growth, leadership, and diversity. At Oracle, we offer unparalleled job opportunities in the tech industry, fostering a culture of innovation and continuous improvement.

Work You’ll Do

At Oracle, your work will directly impact the future of technology across industries. As part of our team, you will lead projects that redefine the way businesses operate, leveraging Oracle’s cutting-edge technology solutions. Our commitment to leadership in the tech community means you’ll be working at the forefront of innovation, enhancing your skills through hands-on experience and comprehensive diversity training.

Join Our Dynamic Team

Oracle is not just a technology company; we are a team of dedicated professionals committed to creating a supportive and inclusive environment. Here, every team member’s contribution is valued, and diversity is celebrated. With Oracle, you are not just accepting a job; you are joining a community that promotes personal and professional growth through constant learning and development opportunities.

Innovative Work and Career Advancement

Embrace the chance to do innovative work with Oracle Corporation, where we push the boundaries of what is possible. With over 130,000 dedicated professionals globally, Oracle offers a workplace where innovation and thought leadership thrive. This environment is perfect for those who are driven to explore new ideas and are eager for opportunities to advance their careers.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional looking for your next career challenge or a student seeking a promising internship, Oracle provides a range of opportunities. Explore positions that match your skills and interests in areas such as cloud computing, enterprise software, and business analytics. Our hiring process is designed to find not just the right skills but also the right fit for Oracle’s unique culture.

Benefits and Culture

Oracle is committed to supporting our employees’ life and work ambitions. We offer competitive benefits, including health insurance, retirement plans, and wellness programs, all designed to support your career and well-being. Our culture of empowerment encourages networking and collaboration across teams and geographies, ensuring that innovation and creativity flourish.

Develop Your Skills Through Training and Networking

Prepare for your future with Oracle’s comprehensive training programs. From leadership development to technical skills enhancement, we provide the tools necessary to succeed in your career and stay ahead in the industry. Networking within Oracle’s global community will also open doors to collaborative opportunities and career advancement.

Stay Connected with Oracle Careers

Keep up to date with the latest from Oracle Corporation by following our careers blog. Gain insights from the experts and learn about new job openings as they become available. Personalize your job search and stay informed about Oracle’s career events and professional development opportunities.

Join Oracle Corporation—Where Careers Grow

At Oracle, we believe in nurturing the potential of our employees. The growth of our company is driven by the individual successes of our team members. We invite you to bring your unique talents to Oracle, join our mission to drive technological innovation, and help shape the future of the digital world.

Search Oracle Jobs

Ready to take the next step in your career? Search for open positions that align with your skills and passions. We are continuously looking for curious, creative, and motivated individuals to join our team. Explore the opportunities and find out how you can contribute to the success of Oracle Corporation.

Oracle Corporation: Leadership, Innovation, Opportunity.

Learn more about Oracle Corporation
Size
143,000 employees
Market Cap
$217.3 billion
Industry
Net Income
$12.8 billion
Founded
1977
5 Year Trend
+2.3%
Revenue
$39.6 billion
NASDAQ

Similar Jobs

More Jobs at Oracle Corporation

More Information Technology Jobs

Find similar Lead Principal Site Reliability Engineer jobs: