We are working on solving some of the most challenging and interesting technology projects around, on a scale unmatched by most. We want people who are passionate about optimizing and troubleshooting data center hardware at mega-scale so our customers don't have to worry! Reporting to the Manager of Infra::Machines::Design, you will support sustaining engineering efforts for the hardware infrastructure of the DigitalOcean server fleet.
The ideal candidate will be eager to face new challenges as DigitalOcean continues to scale its data center footprint and infrastructure cloud capacity and explore new technologies and capabilities to bring to our customers, providing top-tier support for our hardware and firmware in the DO fleet.
What You'll Be Doing:- Act a member of the Sustaining Engineering team in the Infra::Machines::Design Organization
- Support server hardware, cabling, and networking hardware throughout its operational lifecycle
- Monitor the #machines channel and MACHINES JIRA project for issues and drive them to resolution
- Participate in 24/7 on-call rotation with other members of the team
- Act as Tier 2 escalation for Datacenter Operations (DCOPS) and Cloud Operations (CloudOps) regarding hardware and firmware components
- Develop and maintain standards and practices for DigitalOcean hardware operations
- Work closely with the Qualification team, Firmware team, Fleet Lifecycle Engineering team (FLE), Foresight team, and Infrastructure Services team to resolve issues in tooling, firmware packages, hardware components, and other operational concerns
- Help with development of tooling and associated runbooks to address gaps in operational capabilities around hardware and firmware operations
- Coordinate with Ops teams on monitoring thresholds, failure modes and alerting
- Assist in troubleshooting cause of failures and work to prevent them in the future
- Raise the quality bar in the delivery of our cloud infrastructure by identifying industry best practices and working to adopt them
What We'll Expect From You:- Technical Degree (BS Computer Science/Engineering) or equivalent practical experience
- Hands-on experience operating a cloud infrastructure at mid-tier scale or better
- An in-depth understanding of server hardware, firmware, and infrastructure
- Strong knowledge in troubleshooting techniques, Python and BASHExtra points for JTAG debugging / Firmware troubleshooting / Wire sniffing experience
- Clear communication and collaboration across key stakeholders
- An insatiable passion for constant improvement
Compensation Range: *This is a hybrid role
#LI-Hybrid
Application Limit: You may apply to a maximum of 3 positions within any 180-day period. This policy promotes better role-candidate matching and encourages thoughtful applications where your qualifications align most strongly.