Oracle Corporation

Principal Systems Software Engineer - GPU Platform Systems

Oracle Corporation$114K — $234K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years of programming/scripting experience in C, C++, Python, or Bash.
  • Strong background in systems or embedded software development.
  • Expertise in hardware/software integration and computer hardware fundamentals.
  • Proven ability to diagnose complex hardware and software interaction problems.
  • Experience with Linux/Unix development environments.
  • Proficiency in software development practices including design, debugging, and testing.

Responsibilities

  • Design and maintain BMC and platform-management software for GPU and server infrastructure.
  • Develop host management, telemetry, and system diagnostics capabilities.
  • Implement secure firmware-management workflows across multi-vendor platforms.
  • Create low-level systems software using C, C++, and Python.
  • Lead initial board and platform bring-up for new systems.
  • Diagnose integration issues across firmware, hardware, and operating systems.
  • Establish and improve RAS capabilities for GPU and server environments.

Benefits

  • Comprehensive medical, dental, and vision insurance.
  • Short and long-term disability coverage.
  • Life insurance, including AD&D and supplemental options.
  • Flexible Spending Accounts for healthcare and dependents.
  • Generous paid time off, including flexible vacation and 11 paid holidays.
Full Job Description
Job Description

Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve the low-level systems software and platform-management capabilities that power next-generation GPU infrastructure.

This is a hands-on systems engineering role at the intersection of software, firmware, hardware, and large-scale cloud infrastructure. You will work on complex GPU and server platforms, owning critical capabilities spanning BMC/service-processor software, platform management, firmware lifecycle, reliability and serviceability, telemetry, power management, hardware bring-up, and fleet operations.

You will work closely with silicon, firmware, hardware, compute, and fleet engineering teams, as well as technology and manufacturing partners, to bring new platforms from initial hardware enablement through production deployment and ongoing operation at cloud scale.

This role is well suited for an engineer with deep experience in computer systems, embedded or platform firmware, and hardware/software integration who enjoys solving difficult problems that cross traditional engineering boundaries.

Responsibilities

Key Responsibilities:

BMC & Platform-Management Systems
  • Design, develop, and maintain complex BMC, service-processor, and platform-management software for GPU and server infrastructure.
  • Develop capabilities for host management, baseboard management, telemetry, platform monitoring, power control, and system diagnostics.
  • Design and implement secure, reliable firmware-management and update workflows across complex, multi-vendor platforms.
  • Define robust interfaces and integration contracts between platform firmware, hardware components, operating systems, drivers, and higher-level infrastructure services.


Systems & Firmware Development
  • Design and implement low-level systems software and firmware using technologies such as C, C++, Python, and Bash.
  • Develop maintainable software for managing, monitoring, diagnosing, and provisioning server and GPU systems.
  • Build automation and tooling that improves platform provisioning, onboarding, validation, diagnostics, and fleet operations.
  • Develop and debug advanced platform capabilities involving areas such as RAS, telemetry, power management, high-speed I/O, and chipset/SoC services.
  • Conduct design and code reviews and help establish strong engineering practices for maintainability, testing, observability, security, and reliability.


Hardware Bring-Up & Cross-Layer Debugging
  • Play a leading role in initial board, device, and platform bring-up for new GPU and server systems.
  • Diagnose difficult failures spanning hardware, firmware, bootloaders, operating systems, drivers, and platform services.
  • Use hardware and software diagnostic techniques-including logs, schematics, JTAG, logic analyzers, emulators, and platform instrumentation-to isolate root causes.
  • Partner with silicon, board, firmware, and manufacturing teams to validate end-to-end platform behavior and resolve integration issues.
  • Turn complex or recurring failures into durable engineering fixes, improved diagnostics, automation, and preventive controls.


GPU Reliability, Serviceability & Operations
  • Develop reliability, availability, and serviceability (RAS) capabilities for large-scale GPU and server environments.
  • Improve telemetry, fault detection, logging, observability, and diagnostics used to identify and resolve platform issues.
  • Develop and improve power-control and power-capping capabilities and related platform instrumentation.
  • Design systems with fleet-scale reliability, fault tolerance, secure firmware lifecycle, and operational serviceability in mind.
  • Support difficult platform incidents and escalations and help translate field findings into long-term product and engineering improvements.


Technical Leadership
  • Own technically complex and sometimes ambiguous areas from architecture and design through implementation, validation, and deployment.
  • Drive technical decisions and establish clear interfaces across teams responsible for different layers of the platform.
  • Lead deep technical investigations and help teams reach evidence-based root causes for difficult system failures.
  • Raise engineering standards through architecture and design reviews, code reviews, testing practices, automation, and diagnostic tooling.
  • Mentor and provide technical guidance to engineers while remaining actively involved in design, coding, bring-up, and debugging.
  • Collaborate effectively across software, firmware, hardware, silicon, compute, fleet, support, and external partner organizations.


In addition, you have:
  • 6+ years of programming and/or scripting experience, with relevant languages such as C, C++, Python, or Bash.
  • Strong systems-software, embedded-software, or firmware development experience.
  • Experience integrating software or firmware with complex hardware systems.
  • Strong understanding of computer hardware fundamentals and the interaction between hardware, firmware, operating systems, and software.
  • Demonstrated ability to diagnose complex problems that cross hardware and software boundaries.
  • Experience with software development practices including design, implementation, debugging, code review, testing, automation, and quality assurance.
  • Experience working in Linux/Unix-based development or systems environments.
  • Ability to independently own technically complex projects and collaborate across multiple engineering organizations.


Preferred Qualifications:

Experience in one or more of the following areas is highly desirable:
  • BMC, OpenBMC, service processors, or server platform-management technologies.
  • Server, GPU, accelerator, or other complex compute-platform firmware.
  • GPU or server RAS, telemetry, fault management, observability, or serviceability.
  • Board, device, or system bring-up.
  • Hardware debugging using JTAG, logic analyzers, emulators, schematics, or related diagnostic tools.
  • Firmware lifecycle management and secure firmware-update mechanisms.
  • Power management, power control/capping, thermal management, or platform telemetry.
  • Hardware interfaces and low-level communication protocols.
  • CPU, GPU, SoC, ASIC, or FPGA-based systems.
  • ARM, AMD, Intel, NVIDIA, or similarly complex compute platforms.
  • Automation and diagnostics for server provisioning, validation, or fleet operations.
  • Large-scale cloud or data-center infrastructure.
  • Technical leadership, mentoring, architecture/design ownership, and cross-functional engineering coordination.


Location: On-Site | Santa Clara, CA

Qualifications

US: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC4

About Oracle Corporation

Oracle Dyn Global Business Unit is a pioneer in managed DNS and a leader in cloud-based infrastructure that connects users with digital content and experiences across a global internet. Dyn's solution is powered by a global network that drives 40 billion traffic optimization decisions daily for more than 3,500 enterprise customers, including preeminent digital brands such as Netflix, Twitter, Linkedin and CNBC. Adding Dyn's best-in-class DNS and email services extend the Oracle cloud computing platform and provides enterprise customers with a one-stop shop for Infrastructure-as-a-Service (IaaS) and Platform-as-a-Service (PaaS). On January 31, 2017 Oracle completed the acquisition of Dyn, which now operates as an Oracle Infrastructure-as-a-Service (IaaS) global business unit (GBU).

Oracle Corporation Careers

Join Oracle Corporation, a global leader in technology and innovation, and be part of a team that values professional growth, leadership, and diversity. At Oracle, we offer unparalleled job opportunities in the tech industry, fostering a culture of innovation and continuous improvement.

Work You’ll Do

At Oracle, your work will directly impact the future of technology across industries. As part of our team, you will lead projects that redefine the way businesses operate, leveraging Oracle’s cutting-edge technology solutions. Our commitment to leadership in the tech community means you’ll be working at the forefront of innovation, enhancing your skills through hands-on experience and comprehensive diversity training.

Join Our Dynamic Team

Oracle is not just a technology company; we are a team of dedicated professionals committed to creating a supportive and inclusive environment. Here, every team member’s contribution is valued, and diversity is celebrated. With Oracle, you are not just accepting a job; you are joining a community that promotes personal and professional growth through constant learning and development opportunities.

Innovative Work and Career Advancement

Embrace the chance to do innovative work with Oracle Corporation, where we push the boundaries of what is possible. With over 130,000 dedicated professionals globally, Oracle offers a workplace where innovation and thought leadership thrive. This environment is perfect for those who are driven to explore new ideas and are eager for opportunities to advance their careers.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional looking for your next career challenge or a student seeking a promising internship, Oracle provides a range of opportunities. Explore positions that match your skills and interests in areas such as cloud computing, enterprise software, and business analytics. Our hiring process is designed to find not just the right skills but also the right fit for Oracle’s unique culture.

Benefits and Culture

Oracle is committed to supporting our employees’ life and work ambitions. We offer competitive benefits, including health insurance, retirement plans, and wellness programs, all designed to support your career and well-being. Our culture of empowerment encourages networking and collaboration across teams and geographies, ensuring that innovation and creativity flourish.

Develop Your Skills Through Training and Networking

Prepare for your future with Oracle’s comprehensive training programs. From leadership development to technical skills enhancement, we provide the tools necessary to succeed in your career and stay ahead in the industry. Networking within Oracle’s global community will also open doors to collaborative opportunities and career advancement.

Stay Connected with Oracle Careers

Keep up to date with the latest from Oracle Corporation by following our careers blog. Gain insights from the experts and learn about new job openings as they become available. Personalize your job search and stay informed about Oracle’s career events and professional development opportunities.

Join Oracle Corporation—Where Careers Grow

At Oracle, we believe in nurturing the potential of our employees. The growth of our company is driven by the individual successes of our team members. We invite you to bring your unique talents to Oracle, join our mission to drive technological innovation, and help shape the future of the digital world.

Search Oracle Jobs

Ready to take the next step in your career? Search for open positions that align with your skills and passions. We are continuously looking for curious, creative, and motivated individuals to join our team. Explore the opportunities and find out how you can contribute to the success of Oracle Corporation.

Oracle Corporation: Leadership, Innovation, Opportunity.

Learn more about Oracle Corporation
Size
143,000 employees
Market Cap
$217.3 billion
Industry
Net Income
$12.8 billion
Founded
1977
5 Year Trend
+2.3%
Revenue
$39.6 billion
NASDAQ

Similar Jobs

More Jobs at Oracle Corporation

More Enterprise Technology Jobs

Find similar Principal Systems Software Engineer - GPU Platform Systems jobs: