THE ROLEDrive the technical resolution and stability of Everpure's highest-stakes cloud storage deployments for top-tier hyperscale partners. As a Senior Software Engineer on the Hyperscale Product Escalations team (known internally as a Forensics Engineer), you will serve as the technical bridge between field telemetry and core engineering, diagnosing high-complexity distributed systems failures and creating automation that anticipates issues before they impact customer workloads. You will directly influence customer trust, enable multi-million dollar account success, and help architect the future of hyperscale storage solutions built for AI-driven infrastructure.
WHAT YOU'LL DO- Lead Root-Cause Investigations: Analyze complex system behavior, memory dumps, field telemetry, and low-level code across platform, firmware, and software layers to isolate and resolve critical failures in large-scale distributed systems.
- Build System Health Automation: Design and deploy automated diagnostics, predictive health-monitoring services, and AI-assisted workflow tools that reduce manual triage time and prevent recurring failure patterns across growing global fleets.
- Direct Hyperscaler Engineering Collaboration: Partner directly with technical teams at hyperscale customer organizations to solve deep integration challenges, ensuring platform reliability and unlocking major expansion opportunities.
- Engineer Fleet-Wide Reliability Improvements: File, prioritize, and drive long-term engineering fixes from escalation findings back into core product roadmaps to continuously elevate system stability.
WHAT YOU BRING- Systems & Storage Expertise: Deep hands-on experience in Linux systems engineering, platforms, or low-level firmware, along with a strong grasp of storage technologies (such as SSDs, NVMe, NAND, or distributed storage architectures).
- Advanced Debugging & Software Engineering: Proven mastery of troubleshooting complex distributed systems using languages like Python, Go, or C++, paired with the ability to build automated tools for log analysis, telemetry, and system diagnostics.
- Technical Communication & Problem-Solving: Ability to deconstruct intricate distributed systems failure modes into clear technical action plans, facilitating seamless collaboration with cross-functional SMEs and external engineering leadership.
- Location & Work Environment: We are primarily an in-office environment and therefore, you will be expected to work from the Santa Clara office in compliance with Everpure's policies, unless you are on PTO, or work travel, or other approved leave.
#LI-ONSITE
Salary ranges are determined based on role, level and location. For positions open to candidates in multiple geographical locations, the base salary range is reflective of the labor market across the applicable locations.
This role may be eligible for incentive pay and/or equity.
There is no application deadline and we accept applications on an ongoing basis until the job is filled.
The annual base salary range is:
$149,000-$224,000 USD
WHAT YOU CAN EXPECT FROM US:- Innovation: We celebrate those who think critically, like a challenge, and aspire to be trailblazers.
- Growth: We give you the space and support to grow along with us and to contribute to something meaningful. We have been named Fortune's Best Workplaces in Technology™, Fortune's Best Workplaces in the Bay Area™, and certified as a Great Place to Work®!
- Team: We build each other up and set aside ego for the greater good.
And because we understand the value of bringing your full and best self to work, we offer a variety of perks to manage a healthy balance, including flexible time off, wellness resources, and company-sponsored team events. Check out purebenefits.com for more information.