Staff, Reliability Engineer

Tenstorrent$110K — $130K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in reliability engineering, preferably in high-performance computing or AI hardware.
  • Proficient in statistical analysis methods like HALT, HASS, ALT, and FMEA.
  • Experience addressing complex technical issues in thermal environments and communicating risks to leadership.
  • Strong collaboration skills to integrate mechanical, electrical, and software teams under tight deadlines.

Responsibilities

  • Define the reliability strategy for next-gen AI computing systems.
  • Lead root-cause analysis on failures and implement corrective measures across engineering.
  • Validate and ensure the reliability of advanced cooling systems for AI hardware.
  • Act as the main liaison with manufacturing partners regarding reliability standards.
  • Mentor junior engineers and facilitate high-level design reviews to enhance team capabilities.

Benefits

  • Hybrid work model with flexibility in Toronto, Canada location.
  • Opportunities for skill development in cutting-edge AI technologies.
  • Collaboration with top-tier experts across various engineering disciplines.
  • Exposure to various stages of product development from architecture to deployment.
Full Job Description
Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI. This role ishybrid, based out of Toronto, Canada. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are - You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems. - You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory. - You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership. - You're good at bringing people together, mechanical, electrical, thermal, software, and System Dev & Compliance Validation teams, especially when timelines are tight. - You hold a Bachelor's or Master's in Mechanical Engineering, Electrical Engineering, Reliability Engineering, or a related field. What We Need - Someone to set the reliability strategy for our next-generation AI computing systems, with MTBF modeling as a core piece. - A strong problem-solver who can lead root-cause investigations and drive fixes across engineering and the supply chain. - Someone who summarizes findings and feeds them back to the Systems Engineering design team, reliability as an ongoing loop, not a one-time check. - A close partner to System Dev & Compliance Validation, helping set hardware up for success ahead of formal testing and certification. - Willingness to travel to third-party test facilities for HALT/HASS, to preplan, oversee testing, resolve DUT issues, and assess design risk in person. What You Will Learn - How to build a reliability strategy from scratch for hardware that's pushing the boundaries of AI computing. - How to build predictive models, including MTBF frameworks and accelerated life testing. - How to turn test findings into design improvements through close collaboration with Systems Engineering. - How reliability work sets the stage for validation and certification success. - What it takes to run hands-on testing at manufacturing and test partner sites, including in Taiwan. Tenstorrent offers a highly competitive compensation package and benefits.

About Tenstorrent

Tenstorrent is a semiconductor company that designs and develops computer processors for artificial intelligence and machine learning applications. The company's processors are designed to be energy-efficient and scalable, and are used by a range of businesses, from small startups to large enterprises. Tenstorrent was founded in 2016 and is headquartered in Toronto, Ontario.
Learn more about Tenstorrent
Size
51 employees
Industry
Founded
2016

Similar Jobs

More Jobs at Tenstorrent

More Enterprise Technology Jobs

Find similar Staff, Reliability Engineer jobs: