Oracle Corporation

Lead Principal Site Reliability Engineer - CSS

Oracle Corporation$96K — $264K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, IT, Engineering, or equivalent experience.
  • 8+ years in Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Engineering.
  • Experience with production environments requiring high availability.
  • Familiarity with Oracle Cloud Infrastructure or other major cloud providers.
  • Proficiency with Kubernetes, Docker, and related technologies.
  • Expertise in Infrastructure as Code and CI/CD tools (e.g., Terraform, Jenkins).
  • Strong scripting capabilities in Python, Bash, or PowerShell.

Responsibilities

  • Ensure high availability and performance of applications and cloud infrastructure.
  • Design automation solutions to enhance operations efficiency.
  • Develop Infrastructure as Code using Terraform and automation scripts.
  • Build and manage CI/CD pipelines for deployments.
  • Monitor systems with observability tools to ensure reliability.
  • Collaborate with development teams on improving application resilience.
  • Conduct root cause analysis for incidents and improve processes.

Benefits

  • Comprehensive medical, dental, and vision insurance.
  • 401(k) Savings and Investment Plan with company match.
  • Generous paid time off policy including flexible vacation days and holidays.
  • Paid parental leave and adoption assistance.
  • Employee Stock Purchase Plan and other financial planning resources.
Full Job Description
Job Description

We are seeking a highly motivated Site Reliability Engineer (SRE) to support a strategic customer's cloud platform and mission-critical applications. The SRE will be responsible for ensuring high availability, operational excellence, automation, and continuous improvement of production environments.

The successful candidate will partner with application development, platform engineering, cloud infrastructure, networking, cybersecurity, and operations teams to build resilient, secure, and highly automated systems. This role emphasizes reliability engineering, observability, incident response, capacity planning, and proactive operational improvements.

Serves as a consultant and leads the design and architecture of infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams to develop reliable and scalable infrastructures. Oversees incident response and/or maintenance tasks. Provides strategic, future-oriented health and performance reporting. Contributes to strategies for automation and reviews the development and implementation of automation. Leads the implementation of innovative tools and provides expertise in site reliability trends.

Responsibilities

Key Responsibilities
  • Maintain high availability, reliability, and performance of enterprise applications and cloud infrastructure.
  • Design and implement automation to reduce manual operational effort and improve deployment consistency.
  • Develop Infrastructure as Code (IaC) using Terraform and automation scripts using Python, Bash, or similar languages.
  • Build and maintain CI/CD pipelines to support automated deployments and release management.
  • Monitor applications and infrastructure using observability platforms, including metrics, logs, traces, and alerting.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Participate in on-call rotations and respond to production incidents with urgency and professionalism.
  • Conduct root cause analysis (RCA) and implement corrective and preventive actions to eliminate recurring issues.
  • Perform capacity planning, performance tuning, and scalability assessments.
  • Support Kubernetes clusters, containerized workloads, and cloud-native applications.
  • Collaborate with development teams to improve application resiliency, fault tolerance, and operational readiness.
  • Partner with security teams to ensure systems comply with enterprise security and regulatory requirements.
  • Create operational runbooks, documentation, and standard operating procedures.
  • Continuously improve platform reliability through automation, monitoring, and operational best practices.
  • Takes full ownership of forecasting infrastructure demands and strategically responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads, anticipating resource gaps.
  • Seeks opportunities for prototyping and encourages the implementation of prototyping initiatives (e.g., testing new applications or infrastructures, assisting in onboarding), driving new approaches.
  • Identifies and projects resource gaps using cost calculators, identifying alternative solutions to lower costs, when necessary.
  • Provides guidance to others when monitoring services, maintains up-to-date knowledge of their performance, and meticulously documents their condition.
  • Oversees root cause analyses for incidents and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery), ensuring efficient execution and preventing incident reoccurrence.
  • Provides strategic, future-oriented health and performance reporting, and anticipates actions needed based on trends in data.
  • Provides expert-level release notes, communication, and/or guidance on the scale, capacity, security, performance attributes, and requirements of services and technology to customers, cross-functional teams, leadership, and external stakeholders.
  • Provides leadership in the on-call shifts.
  • Leads the resolution of multifaceted technical issues spanning multiple services, functions, and customers, collaborating with cross-functional teams and leveraging advanced investigation and debugging techniques to ensure the achievement of SLOs (service level objectives).
  • Provides input on strategic initiatives to address opportunities to improve performance bottlenecks and deployments, maximizing resource utilization, cost efficiency, speed, and scalability across their line of business.
  • Provides expertise in site reliability trends, guiding the creation and sharing of insights and best practices to shape the future of building, testing, deploying, and running services.


Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • 8+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Engineering.
  • Experience supporting production environments with strict availability requirements.
  • Hands-on experience with Oracle Cloud Infrastructure (OCI) or another major cloud provider (AWS, Azure, or GCP).
  • Experience with Kubernetes, Docker, and container orchestration platforms.
  • Proficiency in Infrastructure as Code (Terraform preferred).
  • Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
  • Strong scripting skills in Python, Bash, or PowerShell.
  • Experience with Linux system administration.
  • Strong understanding of networking concepts including DRG, DNS, load balancing, routing, APIs, and Endpoints.
  • Experience with Observability tools such as Prometheus, Grafana, ELK/OpenSearch, Splunk, Datadog, New Relic, or OCI Native Observability services.
  • Strong troubleshooting and problem-solving skills.
  • Experience supporting financial services or other highly regulated industries.
  • Knowledge of disaster recovery, backup strategies, and business continuity.


Desired Skills
  • Site Reliability Engineering (SRE)
  • Oracle Cloud Infrastructure (OCI)
  • Kubernetes
  • Docker
  • Terraform
  • Python
  • Bash
  • Linux Administration
  • CI/CD Pipelines
  • Git
  • Monitoring & Observability
  • Incident Management
  • Root Cause Analysis (RCA)
  • Automation
  • Capacity Planning
  • Performance Tuning
  • Networking
  • Cloud Security
  • High Availability
  • Disaster Recovery
  • DevOps
  • Agile Methodologies


Success Measures
  • Achieve and maintain service availability and reliability targets.
  • Reduce mean time to detect (MTTD) and mean time to recover (MTTR).
  • Increase deployment frequency while minimizing change failure rates.
  • Reduce operational toil through automation and self-service capabilities.
  • Improve platform observability and proactive issue detection.
  • Meet defined SLOs and error budget objectives.
  • Deliver resilient, secure, and scalable cloud platforms that support NFCU's critical business services.


This role is ideal for an engineer who is passionate about automation, operational excellence, and building highly reliable cloud platforms. The successful candidate will play a key role in ensuring NFCU's production environments remain secure, resilient, and capable of supporting mission-critical financial services at scale.

#LI-JC1

Qualifications

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $96,300 to $264,100 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC5

About Oracle Corporation

Oracle Dyn Global Business Unit is a pioneer in managed DNS and a leader in cloud-based infrastructure that connects users with digital content and experiences across a global internet. Dyn's solution is powered by a global network that drives 40 billion traffic optimization decisions daily for more than 3,500 enterprise customers, including preeminent digital brands such as Netflix, Twitter, Linkedin and CNBC. Adding Dyn's best-in-class DNS and email services extend the Oracle cloud computing platform and provides enterprise customers with a one-stop shop for Infrastructure-as-a-Service (IaaS) and Platform-as-a-Service (PaaS). On January 31, 2017 Oracle completed the acquisition of Dyn, which now operates as an Oracle Infrastructure-as-a-Service (IaaS) global business unit (GBU).

Oracle Corporation Careers

Join Oracle Corporation, a global leader in technology and innovation, and be part of a team that values professional growth, leadership, and diversity. At Oracle, we offer unparalleled job opportunities in the tech industry, fostering a culture of innovation and continuous improvement.

Work You’ll Do

At Oracle, your work will directly impact the future of technology across industries. As part of our team, you will lead projects that redefine the way businesses operate, leveraging Oracle’s cutting-edge technology solutions. Our commitment to leadership in the tech community means you’ll be working at the forefront of innovation, enhancing your skills through hands-on experience and comprehensive diversity training.

Join Our Dynamic Team

Oracle is not just a technology company; we are a team of dedicated professionals committed to creating a supportive and inclusive environment. Here, every team member’s contribution is valued, and diversity is celebrated. With Oracle, you are not just accepting a job; you are joining a community that promotes personal and professional growth through constant learning and development opportunities.

Innovative Work and Career Advancement

Embrace the chance to do innovative work with Oracle Corporation, where we push the boundaries of what is possible. With over 130,000 dedicated professionals globally, Oracle offers a workplace where innovation and thought leadership thrive. This environment is perfect for those who are driven to explore new ideas and are eager for opportunities to advance their careers.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional looking for your next career challenge or a student seeking a promising internship, Oracle provides a range of opportunities. Explore positions that match your skills and interests in areas such as cloud computing, enterprise software, and business analytics. Our hiring process is designed to find not just the right skills but also the right fit for Oracle’s unique culture.

Benefits and Culture

Oracle is committed to supporting our employees’ life and work ambitions. We offer competitive benefits, including health insurance, retirement plans, and wellness programs, all designed to support your career and well-being. Our culture of empowerment encourages networking and collaboration across teams and geographies, ensuring that innovation and creativity flourish.

Develop Your Skills Through Training and Networking

Prepare for your future with Oracle’s comprehensive training programs. From leadership development to technical skills enhancement, we provide the tools necessary to succeed in your career and stay ahead in the industry. Networking within Oracle’s global community will also open doors to collaborative opportunities and career advancement.

Stay Connected with Oracle Careers

Keep up to date with the latest from Oracle Corporation by following our careers blog. Gain insights from the experts and learn about new job openings as they become available. Personalize your job search and stay informed about Oracle’s career events and professional development opportunities.

Join Oracle Corporation—Where Careers Grow

At Oracle, we believe in nurturing the potential of our employees. The growth of our company is driven by the individual successes of our team members. We invite you to bring your unique talents to Oracle, join our mission to drive technological innovation, and help shape the future of the digital world.

Search Oracle Jobs

Ready to take the next step in your career? Search for open positions that align with your skills and passions. We are continuously looking for curious, creative, and motivated individuals to join our team. Explore the opportunities and find out how you can contribute to the success of Oracle Corporation.

Oracle Corporation: Leadership, Innovation, Opportunity.

Learn more about Oracle Corporation
Size
143,000 employees
Market Cap
$217.3 billion
Industry
Net Income
$12.8 billion
Founded
1977
5 Year Trend
+2.3%
Revenue
$39.6 billion
NASDAQ

Similar Jobs

More Jobs at Oracle Corporation

More Information Technology Jobs

Find similar Lead Principal Site Reliability Engineer - CSS jobs: