Principal Hardware Diagnostics Engineer

Graphcore

$130K — $180K *
Telecommunications & Hardware
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or related discipline.
  • Strong software engineering experience in Python, C++, or C#.
  • Experience developing diagnostics or monitoring systems for hardware platforms.
  • Experience working with distributed systems or cloud infrastructure.
  • Strong knowledge of Linux environments and system-level diagnostics tools.
  • Experience collaborating with CM/ODM partners on manufacturing diagnostics and fault isolation.
  • Strong analytical and debugging skills.
  • Excellent communication and collaboration abilities.

Responsibilities

  • Design and develop automated hardware diagnostics solutions for blade-level servers and rack-scale AI systems.
  • Architect and implement diagnostic agents, monitoring tools, and analytics frameworks to track hardware telemetry.
  • Collaborate with hardware teams to integrate low-level diagnostic modules into monitoring systems.
  • Develop diagnostics tools capable of detecting hardware health conditions and isolating failures.
  • Create diagnostic modules used for internal validation and production data center operations.
  • Provide detailed hardware fault information to system engineers to accelerate troubleshooting.
  • Define remediation workflows and insights for hardware fault scenarios across nodes and clusters.
  • Collaborate with firmware, networking, and cloud platform teams to integrate diagnostics across the system stack.

Benefits

  • Work with a globally recognized leader in AI computing systems.
  • Collaborative team environment integrating hardware and software innovations.
  • Opportunity to influence diagnostics solutions on advanced semiconductors and data center hardware.
  • Focus on both innovation and efficiency for broader AI adoption.
  • Engagement with cutting-edge AI infrastructure that supports a transformative technology.
Full Job Description
Job Summary

We are seeking an experienced Principal Hardware Diagnostics Engineer to design and develop diagnostics software used to monitor hardware health and diagnose system-level issues across Graphcore's AI infrastructure platforms.

This role focuses on building diagnostics agents, tools, and analytics frameworks that enable engineers and automation systems to identify, isolate, and resolve hardware issues across blade-level servers and rack-scale clusters.

The Team

Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption.

The Systems Engineering and Platform Validation team ensures Graphcore's AI compute platforms are reliable, diagnosable, and operationally robust at scale.

The team collaborates with hardware engineering, firmware, cloud infrastructure, and automation teams to develop tools and frameworks that monitor system health, detect hardware failures, and accelerate root-cause analysis across AI clusters.

Responsibilities and Duties
  • Design and develop automated hardware diagnostics solutions for blade-level servers and rack-scale AI systems.
  • Architect and implement diagnostic agents, monitoring tools, and analytics frameworks to track hardware telemetry.
  • Collaborate with hardware teams to integrate low-level diagnostic modules into monitoring systems.
  • Develop diagnostics tools capable of detecting hardware health conditions and isolating failures.
  • Create diagnostic modules used for internal validation and production data center operations.
  • Provide detailed hardware fault information to system engineers to accelerate troubleshooting.
  • Define remediation workflows and insights for hardware fault scenarios across nodes and clusters.
  • Collaborate with firmware, networking, and cloud platform teams to integrate diagnostics across the system stack.

Candidate Profile

Essential
  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or related discipline.
  • Strong software engineering experience in Python, C++, or C#.
  • Experience developing diagnostics or monitoring systems for hardware platforms.
  • Experience working with distributed systems or cloud infrastructure.
  • Strong knowledge of Linux environments and system-level diagnostics tools.
  • Experience collaborating with CM/ODM partners on manufacturing diagnostics and fault isolation.
  • Strong analytical and debugging skills.
  • Excellent communication and collaboration abilities.

Desirable
  • Experience working with AI hardware platforms or accelerator-based computing systems.
  • Familiarity with hyperscale data center infrastructure.
  • Experience building cluster-level monitoring or diagnostics systems.
  • Experience interacting with internal or external customers during diagnostics solution development.

Similar Jobs

More Jobs at Graphcore

More Telecommunications & Hardware Jobs

Find similar Principal Hardware Diagnostics Engineer jobs: