Job DescriptionThe role will provide technical leadership for the architecture, design, and evolution of OCI's high-performance virtual networking dataplane. This role is responsible for shaping the systems and engineering direction behind foundational networking capabilities, including packet processing, routing, network address translation, virtual network gateways, traffic management, and security-policy enforcement in a highly scalable, multi-tenant cloud environment.
The engineer will define and drive the architecture for dataplane services that must deliver high throughput, low latency, strong tenant isolation, fault tolerance, and operational simplicity at cloud scale. They will make critical technical decisions across C/C++ systems design, concurrency, memory management, performance optimization, Linux networking, distributed control-plane integration, and fast dataplane technologies such as DPDK, VirtIO, eBPF/XDP, SR-IOV, SmartNICs, and userspace packet-processing frameworks.
This role leads complex initiatives from early technical strategy through architecture, implementation, testing, rollout, and long-term operation. The Principal Engineer will establish design patterns and technical standards for safe, observable, recoverable services; guide the design of resilient mechanisms such as health checks, failover, graceful draining, rollback, and recovery; and ensure features can withstand failures, in-service updates, traffic spikes, and evolving customer requirements without disruption.
The Principal Engineer will also influence how the organization measures and operates dataplane health. They will drive investments in metrics, telemetry, dashboards, alarms, canaries, diagnostic tooling, automated validation, and incident learning. They will lead investigation of the most complex performance and production issues, using system and network-level evidence to identify root causes and create durable improvements.
Success in this role requires broad technical depth as well as strong cross-functional leadership. The Principal Engineer will partner with control-plane, infrastructure, security, hardware, operations, and product teams to align architecture and execution across service boundaries. They will mentor engineers, raise the quality of design and code reviews, clarify difficult technical tradeoffs, and create reusable engineering practices that improve the reliability, performance, security, and delivery velocity of OCI virtual networking.
ResponsibilitiesKey ResponsibilitiesSystem Design and ArchitectureDataplane Scalability and Performance- Design and implement major C and C++ features for OCI's high-performance virtual networking dataplane, including routing, NAT, network gateways, and security-policy enforcement.
- Own complex packet-processing components that support scalable, multi-tenant cloud networking.
- Drive performance improvements across CPU utilization, memory use, concurrency, latency, throughput, and packet-processing efficiency.
- Apply and evolve fast dataplane technologies such as DPDK, VirtIO, Linux networking, eBPF/XDP, SR-IOV, SmartNICs, and userspace packet-processing frameworks.
- Define and execute performance, scale, stress, and load tests; use results to identify bottlenecks and deliver measurable improvements.
- Lead system-level design and code reviews, ensuring solutions meet expectations for performance, correctness, reliability, and maintainability.
Reliability and Availability- Design and build fault-tolerant gateway and dataplane components that remain available through host failures, network disruptions, and in-service software updates.
- Implement resilient patterns including health checks, graceful draining, retries, timeouts, failover, rollback, and recovery.
- Develop fault-injection, brownout, failure-recovery, and availability tests for critical network paths.
- Drive safe change management through automated validation, staged rollouts, monitoring, rollback, and recovery procedures.
Observability and Operational Excellence- Define and improve metrics, alarms, dashboards, telemetry, canaries, and health checks for dataplane services.
- Independently diagnose and resolve complex Linux, networking, performance, and production issues using logs, metrics, traces, packet captures, profilers, and core dumps.
- Build runbooks, diagnostic tools, and operational automation that reduce time to detect, mitigate, and recover from incidents.
- Participate in on-call, lead incident investigations for owned components, and deliver durable corrective actions.
Security and Multi-Tenant Isolation- Design and implement features that enforce strong tenant isolation, secure traffic handling, and policy enforcement.
- Partner with security teams to identify, prioritize, and remediate vulnerabilities and operational risks.
- Ensure designs, documentation, and operational practices meet applicable security, compliance, and change-management requirements.
Automation and Change Management- Build and maintain automation and infrastructure-as-code for service development, deployment, monitoring, and operations.
- Improve deployment safety through automated testing, pre-production validation, progressive rollout, and rollback guardrails.
- Own major platform end to end: design, implementation, testing, deployment, production validation, and ongoing support.
Core ResponsibilitiesPlanning and Execution- Independently plan and execute complex work, identifying risks, dependencies, and tradeoffs early.
- Manage priorities across multiple initiatives and communicate progress, decisions, and blockers clearly.
- Break ambiguous technical problems into deliverable milestones and drive them to completion.
Collaboration and Technical Leadership- Collaborate across networking, control-plane, infrastructure, security, hardware, and operations teams to deliver shared outcomes.
- Lead technical discussions for owned areas and provide thoughtful design and code-review feedback.
- Mentor less-experienced engineers in systems programming, debugging, operational practices, and reliable service ownership.
- Document technical decisions and contribute to team standards, reusable patterns, and engineering knowledge.
Problem Solving and Continuous Improvement- Resolve complex, cross-layer issues by analyzing evidence from systems, networking, and production telemetry.
- Make sound tradeoffs among performance, reliability, security, maintainability, and delivery speed.
- Identify and lead practical improvements to development, testing, deployment, observability, and operational workflows.
- Stay current with C/C++, cloud networking, Linux, and high-performance dataplane technologies; share relevant knowledge with the team.
QualificationsDisclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.Range and benefit information provided in this posting are specific to the stated locations onlyUS: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.