Staff, Reliability Engineer

Tenstorrent$120K — $140K *
Technical Services
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in reliability engineering, preferably in high-performance computing or AI hardware.
  • Proficient in statistical analysis methods like HALT, HASS, ALT, and FMEA.
  • Experience addressing complex technical issues in thermal environments and communicating risks to leadership.
  • Strong collaboration skills to integrate mechanical, electrical, and software teams under tight deadlines.

Responsibilities

  • Define the reliability strategy for next-gen AI computing systems.
  • Lead root-cause analysis on failures and implement corrective measures across engineering.
  • Validate and ensure the reliability of advanced cooling systems for AI hardware.
  • Act as the main liaison with manufacturing partners regarding reliability standards.
  • Mentor junior engineers and facilitate high-level design reviews to enhance team capabilities.

Benefits

  • Hybrid work model with flexibility in Toronto, Canada location.
  • Opportunities for skill development in cutting-edge AI technologies.
  • Collaboration with top-tier experts across various engineering disciplines.
  • Exposure to various stages of product development from architecture to deployment.
Full Job Description
Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI.

This role ishybrid, based out of Toronto, Canada.

We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.

Who You Are
  • You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems.
  • You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory.
  • You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership.
  • You're good at bringing people together, mechanical, electrical, thermal, and software teams, especially when timelines are tight.

What We Need
  • Someone to set the reliability strategy for our next-generation AI computing systems, across both data center and workstation products.
  • A strong problem-solver who can lead root-cause investigations on failures and follow through with fixes across engineering and the supply chain.
  • Someone with real experience validating advanced cooling systems, vapor chambers, heat pipes, direct-to-chip liquid cooling.
  • A steady point of contact for our manufacturing partners and suppliers, making sure reliability standards hold up through the NPI process.
  • Someone who can mentor other engineers and lead design reviews that raise the bar for the team.

What You Will Learn
  • How to shape the reliability strategy for some of the industry's most advanced AI computing platforms built around Tenstorrent's cutting-edge silicon.
  • How reliability engineering influences every stage of product development, from architecture and validation through manufacturing and field deployment.
  • How to collaborate with world-class experts across silicon, thermal, mechanical, electrical, and software engineering to solve large-scale system challenges.
  • How to validate emerging cooling technologies, including advanced air cooling and direct-to-chip liquid cooling, for next-generation AI infrastructure.
  • How to influence the future of AI hardware by helping build highly reliable systems that power tomorrow's largest AI workloads.


Tenstorrent offers a highly competitive compensation package and benefits.

About Tenstorrent

Tenstorrent is a semiconductor company that designs and develops computer processors for artificial intelligence and machine learning applications. The company's processors are designed to be energy-efficient and scalable, and are used by a range of businesses, from small startups to large enterprises. Tenstorrent was founded in 2016 and is headquartered in Toronto, Ontario.
Learn more about Tenstorrent
Size
51 employees
Industry
Founded
2016

Similar Jobs

More Jobs at Tenstorrent

More Technical Services Jobs

Find similar Staff, Reliability Engineer jobs: