Job DescriptionThe Lead Principal Core Infrastructure Engineer mentors engineering teams and leads the architecture of highly scalable, interdependent distributed systems. This role identifies and eliminates performance and scalability bottlenecks for hyperscale workloads, defines scalability requirements with stakeholders, and designs elastic, high-impact systems while driving innovation in data plane platforms.
The engineer architects and oversees fault-tolerant, in-service-upgradable systems, optimizes resilience mechanisms such as load shedding, throttling, and rate limiting, and establishes SLO-aligned standards for durability and availability across dependent services. The role also defines KPIs and advanced telemetry, applies formal verification techniques for complex features, and develops robust replication and synchronization strategies.
Additionally, this position leads the resolution of complex production issues, establishes operational readiness and standard operating procedures, directs incident response and root cause analyses, architects advanced security controls, drives compliance and remediation efforts, and delivers enterprise-scale automation through Infrastructure as Code (IaC) and automated patching, updates, and rollback strategies.
ResponsibilitiesKey Responsibilities- Lead the architecture and design of highly scalable, resilient, and distributed cloud infrastructure systems that support hyperscale workloads.
- Drive performance, scalability, and reliability improvements by identifying bottlenecks and implementing architectural optimizations.
- Design fault-tolerant, highly available systems with automated failover, replication, and disaster recovery capabilities while ensuring compliance with Service Level Objectives (SLOs).
- Define and implement observability standards, including KPIs, telemetry, monitoring, and alerting to maintain system health and operational excellence.
- Apply advanced engineering practices, including formal verification, distributed systems design, and data consistency strategies, to ensure system correctness and availability.
- Lead the diagnosis and resolution of complex production issues, establish operational readiness standards, and drive continuous improvements through incident response and root cause analysis.
- Architect secure, compliant cloud infrastructure by implementing enterprise security controls, remediation strategies, and governance best practices.
- Champion Infrastructure as Code (IaC), automation, and modern deployment practices to enable safe, reliable, and efficient software delivery.
- Partner with cross-functional engineering, product, and business stakeholders to define technical strategy and deliver high-impact initiatives.
- Mentor senior engineers, influence architectural direction, and foster engineering excellence through technical leadership, hiring, and knowledge sharing.
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
- 12+ years of experience designing, building, and operating large-scale distributed systems in cloud or enterprise environments.
- Deep expertise in distributed systems architecture, cloud infrastructure, networking, storage, and compute platforms.
- Strong experience with high availability, fault tolerance, scalability, observability, and performance optimization for mission-critical services.
- Proficiency in one or more programming languages such as Java, Go, C++, Python, or Rust.
- Experience with Infrastructure as Code (Terraform, Ansible, or similar), CI/CD pipelines, and cloud automation.
- Proven experience leading complex technical initiatives and influencing architectural decisions across multiple engineering teams.
- Strong analytical, troubleshooting, and incident management skills, including leading root cause analysis for complex production issues.
- Excellent communication and collaboration skills with the ability to influence technical and business stakeholders.
Preferred Qualifications
- Master's degree or Ph.D. in Computer Science, Computer Engineering, or a related field.
- Experience building hyperscale cloud services or infrastructure platforms.
- Experience with Kubernetes, container orchestration, and cloud-native technologies.
- Knowledge of formal verification methods (such as TLA+) and distributed consensus algorithms.
- Experience with public cloud platforms such as Oracle Cloud Infrastructure (OCI), AWS, Azure, or Google Cloud Platform (GCP).
- Demonstrated technical thought leadership through patents, publications, open-source contributions, conference presentations, or other industry recognition.
- Experience mentoring senior engineers and driving engineering excellence across large, geographically distributed organizations.
QualificationsUS: Hiring Range in USD from: $146,300 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC5