Data Center Operations and Maintenance Engineering LeaderLocation: Hybrid | Bellevue, WA Area
Titles: Senior and Staff (multiple roles available)
Lead the Operations Organization Behind Next-Generation AI InfrastructureThe OpportunityThis is a leadership role responsible for building and scaling the Operations & Maintenance function supporting mission-critical AI infrastructure. Depending on experience, responsibilities may range from leading a regional operations team to defining the long-term operational strategy across multiple facilities.
You will partner closely with Engineering, Infrastructure, Networking, Hardware, and Customer Operations leadership to create a highly reliable, scalable operating model capable of supporting one of the industry's most advanced AI infrastructure platforms.
What You'll Do- Build, lead, mentor, and grow high-performing Operations & Maintenance teams.
- Develop the operational strategy, organizational structure, and execution model supporting large-scale AI infrastructure.
- Establish world-class operational processes, including incident response, change management, problem management, and service reliability.
- Lead major incident management efforts and executive communications during production events.
- Drive operational excellence through proactive monitoring, observability, automation, and continuous improvement initiatives.
- Partner with Engineering to ensure operational readiness for new infrastructure deployments and platform launches.
- Define service level objectives (SLOs), operational KPIs, and reliability metrics across the infrastructure portfolio.
- Build scalable on-call programs, escalation models, runbooks, and operational governance.
- Champion root cause analysis and long-term corrective actions to improve platform resilience.
- Influence infrastructure architecture and operational tooling to improve availability, efficiency, and customer experience.
- Help shape the long-term operations organization as the company expands globally.
What We're Looking For- Experience leading Operations, Site Reliability, Infrastructure Operations, Data Center Operations, or Production Engineering organizations.
- Proven success building or scaling operations teams within cloud infrastructure, hyperscale environments, AI infrastructure, or large distributed systems.
- Deep expertise in production operations, incident management, service reliability, and operational excellence.
- Experience leading cross-functional teams during high-severity production incidents.
- Strong understanding of infrastructure operations across compute, networking, storage, and hardware environments.
- Demonstrated success building operational processes, organizational structure, and scalable support models in high-growth environments.
- Executive-level communication skills with the ability to influence engineering and business leadership.
- Comfortable operating in an early-stage organization where many systems and processes are being built for the first time.
Preferred Qualifications- Experience supporting hyperscale cloud platforms, GPU infrastructure, AI platforms, HPC environments, or large-scale data centers.
- Experience with modern observability and monitoring platforms such as Grafana, Prometheus, Datadog, or similar technologies.
- Familiarity with incident management platforms including PagerDuty, Opsgenie, or equivalent solutions.
- Experience implementing operational maturity frameworks and reliability engineering best practices.
- Track record of building globally distributed operations organizations.
Compensation- Competitive base pay for Bellevue market
- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance
- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.
Location- Hybrid role based in the Bellevue, WA area.
- Approximately three days per week in the office.
- U.S. work authorization is required. Visa sponsorship is not currently available.