Lam Research

Senior Manager, Reliability Engineering & AIOps

Lam Research$137K — $287K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, or equivalent with 8-10 years of experience.
  • Experience in leading a reliability or operations team and establishing technical direction.
  • Proven experience in incident command and managing postmortem processes for major outages.
  • Strong disaster recovery planning skills across multiple platforms (Azure, AWS, GCP).
  • Hands-on experience with incident management tools at scale, such as PagerDuty.
  • Expertise in capacity planning and performance engineering for global production environments.
  • Demonstrated ability to define and manage service level objectives and error budgets.

Responsibilities

  • Lead and develop the reliability engineering team, maintaining hands-on technical involvement.
  • Define and implement a reliability strategy, including service level objectives and an error-budget policy.
  • Establish and manage a global on-call and incident response model with standardized procedures.
  • Oversee the incident management platform, ensuring effective alerting and escalation processes.
  • Act as incident commander for major incidents, managing communication and postmortem analysis.
  • Develop and execute disaster recovery strategies and conduct regular recovery validation exercises.
  • Drive capacity planning and performance engineering to mitigate operational risks across all platforms.

Benefits

  • Hybrid work location models offering flexibility for in-office and remote work.
  • Comprehensive benefits package supporting employees throughout various life phases.
  • Opportunities for personal and professional development within a global team.
Full Job Description
The group you'll be a part of

You will join the Reliability Engineering team within Infrastructure Platform Engineering. The group keeps Lam's global infrastructure estate available and recoverable across Azure, AWS, GCP, compute, storage, network, and high-performance computing, supporting engineering and operations teams in the US, Japan, Singapore, Malaysia, India, and Korea.

The impact you'll make

As Senior Manager of Reliability Engineering & AIOps, you lead the team that keeps critical infrastructure running and proves it is ready for the next failure. In this role, you will directly contribute to the availability of the systems Lam's engineering, manufacturing, and business teams depend on every day, and you will build the automation that makes outages rare, short, and unremarkable.

What you'll do

  • Lead, hire, and develop the reliability engineering team, owning on-call health while staying technically hands-on.
  • Set the reliability strategy: define the service level objective program, publish an error-budget policy, and drive adoption across platform and service teams.
  • Build and run a follow-the-sun on-call and response model across six regions, with clean handoffs and one consistent set of runbooks and severity definitions worldwide.
  • Own the incident management and paging platform end to end, including services, schedules, escalation policies, and routing, configured as code and tuned so alerts fire on real risk rather than noise.
  • Serve as incident commander on major incidents, own executive and stakeholder communications, and lead blameless postmortems with tracked follow-up.
  • Own disaster recovery strategy and execution across Azure, AWS, GCP, and core infrastructure platforms, including service-tier recovery objectives, backup and restore validation, failover readiness, DR certification, runbook governance, and recurring exercises measured against RTO and RPO targets.
  • Lead capacity planning and performance engineering across Azure, AWS, GCP, compute, storage, network, and HPC platforms, using demand forecasting, utilization trends, growth modeling, and automation to prevent capacity risk and reduce manual operational work.
  • Define and drive AI Ops requirements for reliability engineering across Azure, AWS, and GCP, including Microsoft Copilot, Cursor, GitHub Copilot, and LLM-based operational workflows for incident triage, runbook generation, knowledge retrieval, root-cause analysis, and safe remediation recommendations.
  • This is a full-time role on a standard schedule, with participation in a global on-call rotation


Who we're looking for

  • Bachelor's degree in Computer Science, Engineering, or a related field with 10 years of related experience; or a Master's degree with 8 years of experience; or equivalent experience.
  • Experience leading or mentoring a reliability or operations team and setting technical direction.
  • Proven incident command on major outages, plus ownership of a postmortem process.
  • Strong background in disaster recovery planning across Azure, AWS, GCP, and core infrastructure platforms, including restore validation, failover testing, recovery-objective definition, and corrective action tracking after DR exercises or production incidents.
  • Hands-on ownership of an incident management and paging platform at scale, such as PagerDuty.
  • Experience with capacity planning, performance trending, utilization analysis, and infrastructure demand forecasting for globally distributed production environments across Azure, AWS, GCP, and on-premises platforms.
  • Track record of defining and defending service level objectives and error budgets in production.
  • Working depth in observability tooling (Prometheus, Grafana, Loki, Tempo or equivalent), infrastructure as code (Terraform), and Python or Go.
  • Practical experience applying AI-assisted engineering and operations tools such as Microsoft Copilot, Cursor, GitHub Copilot, or enterprise LLM platforms to improve troubleshooting, automation, documentation, and engineering productivity across Azure, AWS, GCP, and hybrid infrastructure, with clear guardrails for security, privacy, auditability, and production safety.


Preferred qualifications

  • Experience running global, follow-the-sun operations across multiple regions and time zones.
  • Capacity and performance engineering at multi-region scale, including Azure, AWS, GCP, high-performance computing, large storage estates, hybrid cloud infrastructure, and proactive capacity governance.
  • Policy as code, progressive delivery, and chaos engineering in practice.
  • Experience building or operating AI and agent-assisted automation in operations, with a clear view of its failure modes.
  • Experience designing or operating AI Ops capabilities across Azure, AWS, GCP, and hybrid environments, including LLM-grounded knowledge bases, agent-assisted incident workflows, prompt and evaluation practices, and supervised automation that can recommend or propose operational changes before execution.


Lam offers a variety of work location models based on the needs of each role. Our hybrid roles combine the benefits of on-site collaboration with colleagues and the flexibility to work remotely and fall into two categories - On-site Flex and Virtual Flex. 'On-site Flex' you'll work 3+ days per week on-site at a Lam or customer/supplier location, with the opportunity to work remotely for the balance of the week. 'Virtual Flex' you'll work 1-2 days per week on-site at a Lam or customer/supplier location, and remotely the rest of the time.

#LI-DM1

Salary

CA San Francisco Bay Area Salary Range for this position: $137,000.00 - $287,000.00.

The above salary range for this position is relevant to applicants that reside or work onsite in the California, San Francisco Bay Area only. Salary offers will depend on factors that include the location you work from, your level, education, training, specific skills, years of experience and comparison to other employees already in this role. Actual salary may vary from salary offered due to numerous factors including but not limited to unpaid time off, unpaid leave, company mandated shutdown, and other relevant factors.

Our Perks and Benefits

At Lam, our people make amazing things possible. That's why we invest in you throughout the phases of your life with a comprehensive set of outstanding benefits.

About Lam Research

Lam Research Corporation (Lam Research) is a supplier of wafer fabrication equipment and services to the worldwide semiconductor industry. Lam Research designs, manufactures, markets, refurbish, and services semiconductor processing equipment used in the fabrication of integrated circuits. The Company’s etch and clean technologies enable customers to build integrated circuits. Its etch systems shape the microscopic conductive and dielectric layers into circuits that define a chip’s final use and function. Its Customer Support Business Group (CSBG) provides products and services to maximize installed equipment performance and operational efficiency. Its customer base includes semiconductor memory, foundry, and integrated device manufacturers (IDMs) that make DRAM, NAND, and logic devices for these products.

Lam Research Careers

Joining Lam Research offers a unique opportunity to enhance your career at a leading global company at the forefront of innovation and technology in the semiconductor industry. Work You’ll Do At Lam Research, we are committed to advancing the technology landscape through continuous innovation and leadership. Our team of professionals is dedicated to pushing the boundaries of what's possible, making us a pivotal player in the semiconductor industry. By joining our team, you will collaborate with some of the brightest minds in the field, contributing to projects that significantly impact future technology. Transform Your Career Lam Research stands at the intersection of technology, innovation, and leadership. We offer job opportunities that challenge you to excel and push your limits. Our culture fosters growth and embraces diversity, ensuring that every team member can thrive professionally and personally. Join our dynamic team and be part of a company known for its groundbreaking work in developing equipment and services that drive the production of virtually every leading-edge chip in the world. Professional Growth and Development We believe in nurturing the professional growth of our employees by providing ample opportunities for career advancement through leadership and diversity training programs. Our commitment to your career is reflected in our robust benefits package, designed to support you and your family’s health, well-being, and financial future. Internship and Employment Opportunities Start your career path at Lam Research with our internship programs, which offer a hands-on experience in the semiconductor industry. Interns work closely with experienced mentors, gaining invaluable skills and knowledge that prepare them for full-time positions within our company. For seasoned professionals, we offer a variety of positions that leverage your skills to contribute to our mission of continuous improvement and innovation. We are always looking for curious, creative, and driven individuals to join our team. Inclusive Culture and Networking Lam Research is dedicated to creating a diverse and inclusive environment where all employees can thrive. Our culture encourages networking and collaboration, allowing you to connect with colleagues and industry leaders who are as passionate about technology and innovation as you are. Applying at Lam Research Ready to take the next step in your career? Explore the job opportunities at Lam Research by visiting our Careers page. Tailor your resume to highlight your relevant experience and skills, and prepare for an interview that could lead to a multitude of rewarding career paths with us. Stay Connected Keep up to date with the latest company news, employment trends, and career tips by joining our community. Subscribe to receive updates that can help you navigate your professional journey at Lam Research. Join Our Team Search open positions that match your skills and interests. We look for passionate, innovative, and solution-driven team players. Start your journey with Lam Research today, where your work isn’t just a job—it’s a pathway to personal and professional fulfillment. SEARCH LAM RESEARCH JOBS Discover the opportunities waiting for you at Lam Research, where we turn today’s innovations into tomorrow’s technologies.
Learn more about Lam Research
Size
14,100 employees
Market Cap
$54.8 billion
Industry
Net Income
$2.9 billion
Founded
2013
5 Year Trend
+16.5%
Revenue
$11.9 billion
NASDAQ

Similar Jobs

More Jobs at Lam Research

More Enterprise Technology Jobs

Find similar Senior Manager, Reliability Engineering & AIOps jobs: