System Software Engineer, Node & Cluster Management

MatX

$160K — $500K+*
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • BS or higher in Computer Science, Electrical Engineering, or equivalent experience with 8+ years in systems software
  • Strong experience in Linux systems development and debugging low-level software
  • Proficient programming skills in C and familiarity with at least one systems language (Go, Rust, C++, Python)
  • Experience designing and implementing HTTP/REST APIs for hardware management
  • Solid understanding of device drivers, telemetry paths, and BMC-managed subsystems
  • Skilled in debugging across multiple software and hardware interfaces
  • Ability to independently create management solutions against new hardware specifications

Responsibilities

  • Design and build the node-level management plane for AI systems with telemetry and control operations
  • Implement cluster management solutions and failover algorithms
  • Develop management CLI tools for diagnostics, firmware updates, and device recovery
  • Collaborate with firmware engineers to unify management across in-band and out-of-band paths
  • Aggregate health data and device inventory for cluster management
  • Engage with low-level systems to debug and prototype components
  • Automate lab system management for efficient provisioning and testing
  • Define software interfaces for management and telemetry components

Benefits

  • 4 weeks PTO plus 12 company holidays and remote work flexibility
  • Company-subsidized medical, dental, and vision insurance for employees and dependents
  • Retirement plans with up to 5% company contributions and life insurance offerings
  • Annual professional development budget of $1500
  • Daily catered team meals and full reimbursement for commuting expenses
  • Monthly stipends for personal perks and cell/internet reimbursements
  • Comprehensive mental health coverage and parental support policies
  • Significant AI resources for productivity enhancement
Full Job Description
MatX is seeking System Software Engineer to join our team as we create best-in-class silicon for high-performance and sustainable GenAI. Successful candidates for these roles will be responsible for delivering performant and functionally accurate silicon for MatX products across compute, memory management. High-speed connectivity and other key technologies.

The MatX host system software team owns everything that makes our AI silicon and systems usable: from Linux kernel drivers up through node and cluster management. The team also co-owns the BMC/OpenBMC firmware stack, with dedicated firmware engineers, so host software and out-of-band management are designed together rather than bolted together. We're looking for self-driven engineers who can take a hardware spec and a register map and just start building - prototype drivers, low-level utilities that talk directly to the chip, daemons, and tooling - with minimal hand-holding. Each engineer on this team has a primary focus area, but ownership of overlapping components is shared, and you should expect (and want) to venture across the stack.

What You'll Do Here
  • Design and build the node-level management plane for MatX's AI systems: expose node health, inventory, telemetry, and control operations through HTTP/REST endpoints (e.g., Redfish-style or custom APIs)
  • Design and implement cluster management solutions and failover algorithms to minimize downtime
  • Build the management CLI utilities that operators and internal engineers use daily - interacting with the on-node management and telemetry daemons to query state, run diagnostics, update firmware, and recover devices
  • Partner with our BMC firmware engineers to present unified management and observability across in-band and out-of-band paths - so operators see one coherent node, whether data comes from the host daemons or the BMC (e.g., unified Redfish-style views, firmware update orchestration across host and BMC, and recovery flows that work even when the host is down)
  • Extend node-level capabilities to cluster level: fleet-wide health aggregation, device inventory, alerting hooks, and integration points for our customers' own fleet-management systems
  • Get hands-on with the low-level stack: you'll regularly need to drop below the API layer - into the telemetry daemon, driver interfaces, or raw device access utilities - to prototype, debug, or unblock yourself
  • Build tooling and automation for managing lab systems during bring-up: provisioning, test orchestration, regression monitoring
  • Define the software contracts between the on-node daemons, the BMC stack, and the management layer - shared-ownership boundaries you'll co-design
  • Debug production-grade issues spanning management APIs, daemons, kernel drivers, BMC firmware, and hardware
  • Help shape what "manageable at scale" means for a new hardware platform, from single node to full rack to cluster

Who You Are
  • BS or higher in Computer Science, Electrical Engineering, or equivalent practical experience, with 8+ years in systems software - this is not a pure web-services role; deep low-level systems experience is required
  • Strong hands-on Linux systems development experience, including low-level userspace software; comfortable reading and debugging kernel driver and daemon code
  • Strong programming skills in C plus a systems language suited to services and tooling (Go, Rust, C++, and/or Python)
  • Experience designing and building HTTP/REST APIs and CLI tools for hardware or infrastructure management
  • Solid understanding of how the pieces underneath your APIs actually work - device drivers, telemetry paths, PCIe device behavior, BMC-managed subsystems - and the instinct to go look when something misbehaves
  • Experienced debugging across API, daemon, kernel, firmware, and hardware boundaries
  • Comfortable working with firmware engineers to align host-side and BMC-side management capabilities behind common interfaces
  • Self-driven and pragmatic: able to stand up a working management endpoint against brand-new hardware with minimal specification

Bonus Points If You Have
  • Experience with Redfish, OpenBMC, gNMI, IPMI, or other datacenter hardware management standards
  • Cluster/fleet management experience for GPU or accelerator infrastructure
  • Experience with hardware bring-up, lab automation, or manufacturing/qualification test infrastructure
  • Familiarity with firmware update orchestration, secure boot, or attestation flows

Compensation

The US base salary for this full-time position is determined based on a variety of factors including role, experience, location, job-related skills, and relevant education and training. Career length is only a guideline for compensation.
  • Early Career - $160,000 - $275,000 + equity
  • Mid Career - $175,000 - $400,000 + equity
  • Senior Career - $250,000 - $600,000 + equity

What We Offer
  • Time off: 4 weeks PTO (accrued) + 12 company Holidays + up to 3 weeks remote work
  • Health: Company-subsidized Medical (Kaiser or Anthem) for employees & dependents, Guardian Dental and Vision insurances for employee & dependents, and life insurance (employee only), plus HSA and FSA offerings via Lively.
  • Financial Wellbeing: Choose from Roth IRA/ 401K (or both) retirement plans with up to 5% company contribution to 401K (even if you don't contribute). Also, 100% company-paid life insurance (up to $300K) and long-term disability insurances.
  • Professional Development: $1500 Professional Development Budget (per year)
  • Team Meals: MatX provides onsite team lunch & dinner Monday - Friday, with your choice of ordering via WeBox, Specialty's or via our reimbursement system
  • Commute on Us: Commute on our company Uber account, or reimburse your train rides. Either way, we pay 100% for your daily commute.
  • MatX E[x]tras: $50/mo to use on the perk you value most
  • Cell & Internet Reimbursement: $35/mo for cellular and $40/mo for wifi
  • Mental Wellbeing: 100% paid mental health benefit via SpringHealth and Guardian EAP.
  • Support to Parents: Up to 12 weeks paid parental leave regardless of path to parenthood, 10 weeks pregnancy disability leave, flexible return-to-work hours, and Benepass reproductive health & parental benefit.
  • AI Resources: Up to $20K/month plus a dedicated internal AI Tooling Team to support your productivity

Similar Jobs

More Jobs at MatX

More Enterprise Technology Jobs

Find similar System Software Engineer, Node & Cluster Management jobs: