Production Systems Engineer, AI Systems

Meta

$150K — $180K *
Technical Services
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or related technical field, or equivalent experience
  • 8+ years in hardware systems engineering, silicon validation, or NPI for AI servers and accelerators
  • Experience in ASIC bring-up, board-level debugging, or system validation in data centers
  • Proven track record of developing test specifications and validation procedures for complex systems
  • Knowledge of high-speed interconnects like PCIe, NVLink, DDR5, or HBM in AI contexts
  • Expertise in troubleshooting across hardware, firmware, and software domains
  • Experience managing large teams across projects

Responsibilities

  • Lead validation strategies for AI and HPC hardware platforms
  • Conduct hands-on bring-up and characterization of AI server systems and components
  • Create and maintain validation test specifications and procedures
  • Investigate and resolve complex system failures in collaboration with engineering teams
  • Track defects and ensure progress towards milestone goals
  • Identify improvements in testing coverage and methodologies
  • Define deployment readiness standards with capacity engineering teams
  • Guide data analysis and reporting for hardware quality trends
  • Communicate validation status and risks to teams and vendors
  • Collaborate on hardware-software interface requirements for AI systems

Benefits

  • Comprehensive health insurance plans
  • 401(k) retirement plan with company matching
  • Generous paid time off and holiday policies
  • Employee wellness and professional development programs
  • Opportunities for innovation in cutting-edge technology fields
Full Job Description
Meta is seeking a Hardware Systems Engineer to support the new product introduction (NPI) of next-generation AI and high-performance computing infrastructure for large-scale data center deployments. In this role, you will work at the intersection of server systems, AI applications, and data center operations, partnering with hardware design, firmware, software, networking, and capacity engineering teams to validate and scale cutting-edge AI hardware systems from early bring-up through production readiness.

Responsibilities

Lead end-to-end system validation strategies for AI and HPC hardware platforms, including AI accelerators, GPU clusters, and high-bandwidth memory subsystems in data center environments
• Drive hands-on bring-up, characterization, and validation of AI server systems and associated components such as PCIe, NVLink, DRAM, and high-speed networking fabrics
• Develop and maintain test specifications, validation procedures, and debug guides tailored to AI infrastructure NPI programs
• Investigate and root-cause complex system failures spanning silicon, firmware, software, and hardware layers in collaboration with cross-functional engineering teams
• Triage and track hardware and firmware defects through their resolution while maintaining forward progress on NPI program milestones
• Identify gaps in test coverage and drive improvements to test methodologies, tooling, and automation frameworks across the NPI lifecycle
• Partner with AI platform and capacity engineering teams to define acceptance criteria and deployment readiness standards for new AI hardware systems
• Guide data collection, analysis, and reporting efforts to surface systemic hardware quality trends and inform go/no-go decisions for production deployment
• Communicate validation status, risk assessments, and technical findings to internal engineering teams and external hardware vendors
• Collaborate with firmware and software teams to define hardware-software interface requirements for telemetry, diagnostics, and remote management of AI infrastructure

Minimum Qualifications
• Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
• 8+ years of experience in hardware systems engineering, silicon validation, firmware validation, or system-level bring-up for AI servers, GPUs, TPUs, or AI accelerator platforms
• Experience in one or more of the following domains: ASIC bring-up and characterization, board-level debug, firmware validation, or large-scale system validation in data center environments
• Experience in developing test specifications, validation procedures, and debug methodologies for complex hardware systems
• Experience with high-speed interconnects or memory subsystems such as PCIe, NVLink, DDR5, or HBM in the context of AI or HPC system validation
• Experience leading root-cause analysis and troubleshooting of system-level failures across hardware, firmware, and software stacks
• Experience analyzing system telemetry and fleet health data to identify reliability trends and drive engineering improvements
• Experience leading projects with large teams

Preferred Qualifications
• Experience with high-speed interconnects and memory subsystems such as PCIe, NVLink, InfiniBand, DDR5, or HBM in the context of AI or HPC infrastructure operations
• Proficiency in scripting or programming languages such as Python for the automation of infrastructure workflows and data analysis
• Understanding of the AI training process
• Familiarity with Linux-based server environments and data center management tooling used in large-scale production operations
• Experience defining hardware-software interface requirements for telemetry, out-of-band management, or remote diagnostics in data center AI systems

Similar Jobs

More Jobs at Meta

More Technical Services Jobs

Find similar Production Systems Engineer, AI Systems jobs: