Oracle Corporation

Principal Core Infrastructure Engineer

Oracle Corporation$114K — $234K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in designing and building distributed systems
  • Proven expertise in cloud scalability, reliability, and security principles
  • Strong background in high-performance computing and network architecture
  • Experience in operational readiness and incident management
  • Knowledge of automation, Infrastructure as Code (IaC), and deployment tooling
  • Solid understanding of resilience mechanisms and fault-tolerant systems
  • Familiarity with data-plane technologies for large-scale data processing.

Responsibilities

  • Lead the design and development of scalable distributed systems.
  • Optimize code for large-scale hyperscale workloads.
  • Architect fault-tolerant systems for high availability.
  • Implement monitoring and alerting for system health.
  • Proactively resolve complex production issues across the system.
  • Develop automation and deployment strategies for infrastructure management.
  • Mentor engineering peers and contribute to technical direction.

Benefits

  • Opportunity for professional growth and skill development in cutting-edge technologies.
  • Collaborative environment with a focus on engineering excellence.
  • Flexibility in roles, allowing team members to shape their contributions.
  • Involvement in innovative projects with potential for broad impact.
  • Commitment to a diverse and inclusive workplace.
Full Job Description
Job Description

We are looking for experienced engineers to join an incubation team responsible for initiating, prototyping, and transitioning new projects with broad impact across OCI.

One of our key long-term initiatives is the development of a new platform designed to power OCI services at cloud scale. The platform encompasses low-level execution runtimes, application and lifecycle management, and high-level change-management workflows, providing OCI developers with the infrastructure and abstractions they need to build and operate services efficiently.

Our goal is simple: enable OCI developers to focus on building innovative services while the platform provides the scalability, reliability, security, and operational capabilities required to run them at cloud scale.

As part of this team, you will challenge existing engineering assumptions, explore new architectural approaches, and apply your expertise in high-performance and reliable systems to help evolve OCI's infrastructure.

What You'll Do

You will lead the development and begin architecting key components of scalable, elastic distributed systems, taking ownership of their performance, reliability, operability, and security.

You will:
  • Design and build distributed systems at cloud scale, defining and enforcing scalability requirements for the components you own.
  • Optimize critical code and data paths for high-throughput, hyperscale workloads, leveraging data-plane platforms for large-scale retrieval, storage, and processing.
  • Architect resilient and fault-tolerant systems using redundancy, replication, failover, and well-defined policies for handling network and infrastructure partitions.
  • Design systems that support in-service upgrades, safe patching, updates, and rollbacks while minimizing customer impact.
  • Apply load shedding, throttling, rate limiting, backpressure, and other resilience mechanisms to maintain service objectives under overload, partial failures, and unreliable network conditions.
  • Define and maintain Service Level Objectives (SLOs), Key Performance Indicators (KPIs), telemetry, dashboards, and proactive alerting to provide clear visibility into system health and performance.
  • Develop sophisticated validation strategies-including fault injection, brownout testing, failure simulation, and resilience testing-to verify system behavior under adverse conditions.
  • Design and improve replication and synchronization mechanisms that preserve correctness, consistency, and durability across distributed components.
  • Proactively investigate and resolve complex production issues, performing deep technical analysis across system boundaries to identify root causes and long-term improvements.
  • Drive operational readiness, ensuring services are observable, supportable, recoverable, and prepared for production at scale.
  • Implement robust security controls and remediation strategies, while maintaining the documentation and processes necessary to satisfy security and compliance requirements.
  • Build Infrastructure as Code (IaC), automation, and deployment tooling that enables repeatable, reliable, and safe infrastructure and service changes.
  • Contribute to architectural direction, technical standards, and engineering best practices while mentoring peers and raising the technical bar across the team.
  • Prototype and evaluate new technologies and architectural approaches, helping transition successful incubations into production systems and broader OCI adoption.


Responsibilities

As a member of the OCI Technical Strategy and Oversight organization, you will help design, build, and operate highly scalable, reliable, and secure distributed systems that support OCI services at hyperscale. You will take ownership of critical components from architecture and implementation through production operations, while contributing to technical direction, engineering excellence, and the growth of the broader team.

System Design & Architecture

Scalability & Performance
  • Lead the design, development, and implementation of components for scalable, elastic distributed systems, supporting both horizontal and vertical scaling as workload demands evolve.
  • Define scalability and performance requirements for owned components and ensure those requirements are reflected throughout design, implementation, testing, and production operation.
  • Optimize critical code paths and system architectures for high-throughput, large-scale data processing and hyperscale workloads.
  • Design systems with elasticity in mind, enabling resources to scale efficiently both up and down based on demand.
  • Leverage distributed state-management technologies and data-plane platforms to support large-scale data retrieval, storage, processing, and coordination.
  • Develop comprehensive performance, scalability, capacity, and load-testing strategies to validate system behavior under expected and extreme workloads.
    Reliability & Resilience
  • Design and build fault-tolerant, highly available systems capable of remaining operational during failures, maintenance, and in-service updates.
  • Implement redundancy, replication, automatic failover, and recovery mechanisms that minimize disruption and protect critical workloads.
  • Design systems to operate predictably during infrastructure and network failures, including partitions and partial dependency failures, while making appropriate tradeoffs among consistency, availability, and partition tolerance.
  • Implement and optimize resilience mechanisms including load shedding, throttling, rate limiting, backpressure, retries, and graceful degradation.
  • Establish appropriate availability and durability expectations through clearly defined Service Level Objectives (SLOs) and use them to guide architectural and operational decisions.
  • Design systems and components to support upgrades and maintenance with minimal or no customer-visible downtime.
    Observability & System Performance
  • Define meaningful Key Performance Indicators (KPIs), telemetry, and health signals to measure system performance, reliability, capacity, and operational health.
  • Build and customize dashboards, telemetry pipelines, monitoring systems, and alerting mechanisms that proactively identify degradation and emerging issues.
  • Use production telemetry and performance data to identify bottlenecks, capacity constraints, and opportunities for architectural improvement.
    Correctness, Durability & Availability
  • Define and implement functional and correctness requirements for complex features, components, and distributed systems.
  • Design sophisticated validation strategies-including fault injection, brownout testing, failure simulation, and resilience testing-to verify system behavior under adverse conditions.
  • Develop data replication and synchronization mechanisms that maintain correctness, integrity, consistency, durability, and availability across distributed components.
  • Identify complex failure modes and incorporate appropriate safeguards into system architecture and implementation.
    Operational Excellence & Incident Management
  • Take a proactive role in diagnosing, debugging, and resolving complex issues across production components and distributed systems.
  • Maintain deep technical expertise in owned systems to support effective troubleshooting, performance optimization, and production operations.
  • Design and implement strategies that enable zero- or minimal-downtime maintenance, reducing or eliminating the need for customer-facing maintenance windows.
  • Ensure systems meet operational-readiness requirements before entering production, including observability, capacity planning, recovery procedures, documentation, and failure handling.
  • Participate in operational support rotations and provide technical leadership during incident response, mitigation, and recovery.
  • Lead or contribute to root cause investigations, identify systemic improvements, and ensure lessons from incidents are incorporated into future designs.
  • Mentor engineers in debugging, incident response, operational practices, and distributed-systems troubleshooting.
    Security & Compliance
  • Design and implement robust security controls for applications and infrastructure operating in multi-tenant cloud environments.
  • Apply appropriate encryption, authentication, authorization, access-control, and data-protection mechanisms.
  • Identify security gaps and execute remediation plans to reduce risk and strengthen system security.
  • Ensure infrastructure and services meet applicable security, compliance, and regulatory requirements.
  • Maintain accurate security and compliance documentation and incorporate security considerations throughout the development lifecycle.
    Automation & Change Management
  • Develop and maintain automation, tooling, and Infrastructure as Code (IaC) to provision, configure, operate, and manage cloud infrastructure reliably at scale.
  • Establish and follow change-management practices for application and infrastructure patching, upgrades, deployments, and rollbacks.
  • Design systems and components that enable these processes to become increasingly automated, repeatable, observable, and safe.
  • Improve deployment and operational tooling to reduce manual intervention, minimize risk, and accelerate recovery when changes do not behave as expected.
    Core Responsibilities
    Planning & Execution
  • Lead and coordinate moderately complex engineering initiatives, managing priorities, dependencies, timelines, and deliverables to ensure successful execution.
  • Provide technical oversight across multiple workstreams while balancing short-term delivery with long-term architectural objectives.
  • Prioritize and delegate work effectively, monitor progress, identify risks early, and adjust execution plans as resources, requirements, or timelines evolve.
  • Drive projects toward completion while maintaining high standards for engineering quality, reliability, security, and operational readiness.
    Collaboration & Partnership
  • Collaborate across engineering teams and organizational boundaries to align technical direction, expectations, dependencies, and shared objectives.
  • Develop a strong understanding of the needs of business leaders, stakeholders, customers, and partner teams to ensure proposed solutions address meaningful requirements.
  • Communicate technical decisions, tradeoffs, risks, and recommendations clearly to both technical and non-technical stakeholders.
  • Foster an inclusive engineering environment by actively seeking diverse perspectives, encouraging constructive discussion, and ensuring team members feel heard and respected.
    Problem Solving & Technical Judgment
  • Analyze complex technical problems using data, system behavior, telemetry, experimentation, and engineering judgment to identify effective solutions.
  • Investigate issues across component and organizational boundaries rather than limiting analysis to individual services.
  • Proactively escalate critical or unresolved issues with a clear assessment of impact, risks, alternatives, and recommended solutions.
  • Document problem-solving approaches, architectural decisions, tradeoffs, and lessons learned to improve organizational knowledge and future decision-making.
    Continuous Learning & Mentorship
  • Continuously expand expertise in distributed systems, cloud infrastructure, reliability engineering, security, automation, and emerging technologies.
  • Stay current with relevant industry trends, technologies, architectural patterns, and engineering best practices.
  • Actively seek and incorporate feedback to strengthen technical and leadership capabilities.
  • Coach and mentor engineers, sharing technical knowledge and helping others develop stronger design, implementation, debugging, and operational skills.
  • Promote knowledge sharing within and across teams.
    Continuous Improvement
  • Identify opportunities to simplify and improve engineering processes, architectures, tools, protocols, and operational workflows.
  • Develop and recommend improvements that increase engineering velocity, reliability, scalability, security, and operational efficiency.
  • Collaborate with partner teams to implement improvements that span organizational or system boundaries.
  • Evaluate the impact of proposed changes on customers, developers, operators, and other stakeholders.
  • Solicit feedback and continuously explore alternative approaches to improve technical and organizational effectiveness.
    Team & Talent Development
  • Contribute to building and strengthening the engineering organization through technical mentorship and knowledge sharing.
  • Participate in candidate interviews, assess technical and problem-solving capabilities, and provide thoughtful hiring recommendations.
  • Help maintain a high engineering bar while supporting the development and success of existing and incoming team members.


Qualifications

About Oracle Corporation

Oracle Dyn Global Business Unit is a pioneer in managed DNS and a leader in cloud-based infrastructure that connects users with digital content and experiences across a global internet. Dyn's solution is powered by a global network that drives 40 billion traffic optimization decisions daily for more than 3,500 enterprise customers, including preeminent digital brands such as Netflix, Twitter, Linkedin and CNBC. Adding Dyn's best-in-class DNS and email services extend the Oracle cloud computing platform and provides enterprise customers with a one-stop shop for Infrastructure-as-a-Service (IaaS) and Platform-as-a-Service (PaaS). On January 31, 2017 Oracle completed the acquisition of Dyn, which now operates as an Oracle Infrastructure-as-a-Service (IaaS) global business unit (GBU).

Oracle Corporation Careers

Join Oracle Corporation, a global leader in technology and innovation, and be part of a team that values professional growth, leadership, and diversity. At Oracle, we offer unparalleled job opportunities in the tech industry, fostering a culture of innovation and continuous improvement.

Work You’ll Do

At Oracle, your work will directly impact the future of technology across industries. As part of our team, you will lead projects that redefine the way businesses operate, leveraging Oracle’s cutting-edge technology solutions. Our commitment to leadership in the tech community means you’ll be working at the forefront of innovation, enhancing your skills through hands-on experience and comprehensive diversity training.

Join Our Dynamic Team

Oracle is not just a technology company; we are a team of dedicated professionals committed to creating a supportive and inclusive environment. Here, every team member’s contribution is valued, and diversity is celebrated. With Oracle, you are not just accepting a job; you are joining a community that promotes personal and professional growth through constant learning and development opportunities.

Innovative Work and Career Advancement

Embrace the chance to do innovative work with Oracle Corporation, where we push the boundaries of what is possible. With over 130,000 dedicated professionals globally, Oracle offers a workplace where innovation and thought leadership thrive. This environment is perfect for those who are driven to explore new ideas and are eager for opportunities to advance their careers.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional looking for your next career challenge or a student seeking a promising internship, Oracle provides a range of opportunities. Explore positions that match your skills and interests in areas such as cloud computing, enterprise software, and business analytics. Our hiring process is designed to find not just the right skills but also the right fit for Oracle’s unique culture.

Benefits and Culture

Oracle is committed to supporting our employees’ life and work ambitions. We offer competitive benefits, including health insurance, retirement plans, and wellness programs, all designed to support your career and well-being. Our culture of empowerment encourages networking and collaboration across teams and geographies, ensuring that innovation and creativity flourish.

Develop Your Skills Through Training and Networking

Prepare for your future with Oracle’s comprehensive training programs. From leadership development to technical skills enhancement, we provide the tools necessary to succeed in your career and stay ahead in the industry. Networking within Oracle’s global community will also open doors to collaborative opportunities and career advancement.

Stay Connected with Oracle Careers

Keep up to date with the latest from Oracle Corporation by following our careers blog. Gain insights from the experts and learn about new job openings as they become available. Personalize your job search and stay informed about Oracle’s career events and professional development opportunities.

Join Oracle Corporation—Where Careers Grow

At Oracle, we believe in nurturing the potential of our employees. The growth of our company is driven by the individual successes of our team members. We invite you to bring your unique talents to Oracle, join our mission to drive technological innovation, and help shape the future of the digital world.

Search Oracle Jobs

Ready to take the next step in your career? Search for open positions that align with your skills and passions. We are continuously looking for curious, creative, and motivated individuals to join our team. Explore the opportunities and find out how you can contribute to the success of Oracle Corporation.

Oracle Corporation: Leadership, Innovation, Opportunity.

Learn more about Oracle Corporation
Size
143,000 employees
Market Cap
$217.3 billion
Industry
Net Income
$12.8 billion
Founded
1977
5 Year Trend
+2.3%
Revenue
$39.6 billion
NASDAQ

Similar Jobs

More Jobs at Oracle Corporation

More Enterprise Technology Jobs

Find similar Principal Core Infrastructure Engineer jobs: