Lam Research

Senior AI Observability engineer

Lam Research$92K — $211K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 12+ years of relevant experience, Bachelor's degree in Computer Science Engineering or related field.
  • Expert-level knowledge of AWS and Azure, including networking topology and serverless architectures.
  • Hands-on experience implementing cloud-native Observability solutions.
  • Proficiency in monitoring tools such as Prometheus, Grafana, and Datadog.
  • Mastery of Terraform/OpenTofu, and Pulumi for Infrastructure-as-Code (IaC).
  • Expert knowledge of OpenTelemetry and W3C Trace Context.
  • Proficient in Go, Python, or Bash for custom automation scripts.

Responsibilities

  • Design and implement pipelines to collect and aggregate telemetry data from cloud-native sources.
  • Configure AI-driven anomaly detection to proactively identify system issues.
  • Collaborate with software teams to ensure auto-instrumentation in CI/CD pipelines.
  • Automate the deployment of dashboards and alerting rules using IaC.
  • Leverage AIOps and distributed tracing to enhance system performance and reduce MTTR.
  • Monitor security event logs to ensure compliance and identify vulnerabilities.
  • Automate infrastructure provisioning using Terraform and other tools.

Benefits

  • Comprehensive benefits to support various life phases.
  • Hybrid work models offering flexibility in remote and on-site work arrangements.
Full Job Description
The group you'll be a part of

We are seeking a hands-on Senior AIOps Reliability Engineer to build the AI-native operations layer for our hybrid enterprise estate. You will design and ship LLM-based agents, retrieval-augmented knowledge pipelines, and machine-learning anomaly detection that find, triage, and remediate incidents across public cloud, private cloud, and on-premises data centers, all resting on strong SRE and network engineering fundamentals.

Our environment spans Azure, AWS, and Google Cloud alongside on-premises data centers, colocation sites, and manufacturing and HPC facilities, so cloud-agnostic design and hybrid network fluency matter more than depth in any single provider. This is a deep individual-contributor role: you will write the code, instrument the telemetry, tune the models, and own the reliability of both the infrastructure and the AI systems operating on it.

The impact you'll make

Join Lam as an IT Engineer, where you'll be at the forefront of designing, analyzing, and implementing applications and systems that form the foundation of our infrastructure. As a crucial member of our IT team, you'll contribute your technical assistance and guidance to projects for various systems and infrastructures. Acting as a technical liaison, you'll address complex business problems with automated systems solutions. Your expertise will be instrumental in driving Lam's commitment to innovation and efficiency.

What you'll do

AI and Agentic Operations
  • Build agentic AI workflows using LLM agents, tool and function calling, and orchestration frameworks such as LangGraph, Semantic Kernel, AutoGen, or the Model Context Protocol, applied to autonomous fault detection, triage, and remediation.
  • Develop the AIOps intelligence layer: time-series anomaly detection, dynamic baselining, alert deduplication and correlation, event clustering, and predictive failure and capacity forecasting across infrastructure, network, and application telemetry from both cloud and on-premises sources.
  • Engineer the retrieval knowledge fabric by chunking, embedding, and indexing runbooks, post-mortems, architecture documents, CMDB and ServiceNow records into a vector store, then tuning retrieval quality against measurable evaluations.
  • Ship AI-assisted incident response: automated summarization, root-cause hypothesis generation, blast-radius analysis, and telemetry-grounded draft post-mortems wired into the paging and ITSM toolchain.
  • Automate remediation safely through event-driven pipelines and configuration-management runbooks invoked by agents, with human-in-the-loop approval gates, scoped least-privilege boundaries, rollback paths, and complete audit trails for every autonomous action.
  • Own AI safety and governance in production: guardrails, prompt-injection defense, hallucination and drift monitoring, PII redaction, and evaluation harnesses that gate every model or prompt change.
  • Run LLMOps and MLOps, covering prompt and model versioning, offline and online evaluation, shadow and A/B testing, inference logging, token cost and latency observability, and CI/CD for every AI component.
  • Instrument AI systems as first-class services with OpenTelemetry GenAI tracing, model SLOs, and quality, cost, and latency dashboards for every agent in production.


Who we're looking for

  • BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Eight or more years in SRE, DevOps, infrastructure, observability, network engineering, or platform engineering, with a track record of shipping production systems yourself.
  • Production experience with LLM applications: prompt engineering, retrieval-augmented generation, embeddings and vector databases, function and tool calling, and agent orchestration.
  • Practical use of machine learning for anomaly detection, forecasting, event correlation, and alert noise reduction on real operational telemetry.
  • Working knowledge of at least one major AI platform and the ability to remain portable across them, including private-network deployment, quota, and cost management.
  • Hands-on production experience with at least one major public cloud and the ability to design portable, cloud-agnostic patterns across the others, covering identity and access boundaries, compute, storage, managed databases, serverless functions, and event services.
  • Solid on-premises infrastructure background across virtualization, storage, data-center operations, and private cloud platforms.
  • Strong networking fundamentals spanning both worlds: TCP/IP, BGP and OSPF, VLANs and overlays such as VXLAN and EVPN, MPLS and SD-WAN, firewalls, load balancers, DNS, and cloud virtual network, VPC, and transit routing constructs. You can read a flow log or a routing table and reason about a failure.


Preferred qualifications

Lam offers a variety of work location models based on the needs of each role. Our hybrid roles combine the benefits of on-site collaboration with colleagues and the flexibility to work remotely and fall into two categories - On-site Flex and Virtual Flex. 'On-site Flex' you'll work 3+ days per week on-site at a Lam or customer/supplier location, with the opportunity to work remotely for the balance of the week. 'Virtual Flex' you'll work 1-2 days per week on-site at a Lam or customer/supplier location, and remotely the rest of the time.

Salary

CA San Francisco Bay Area Salary Range for this position: $92,000.00 - $211,000.00.

The above salary range for this position is relevant to applicants that reside or work onsite in the California, San Francisco Bay Area only. Salary offers will depend on factors that include the location you work from, your level, education, training, specific skills, years of experience and comparison to other employees already in this role. Actual salary may vary from salary offered due to numerous factors including but not limited to unpaid time off, unpaid leave, company mandated shutdown, and other relevant factors.

Our Perks and Benefits

At Lam, our people make amazing things possible. That's why we invest in you throughout the phases of your life with a comprehensive set of outstanding benefits.

About Lam Research

Lam Research Corporation (Lam Research) is a supplier of wafer fabrication equipment and services to the worldwide semiconductor industry. Lam Research designs, manufactures, markets, refurbish, and services semiconductor processing equipment used in the fabrication of integrated circuits. The Company’s etch and clean technologies enable customers to build integrated circuits. Its etch systems shape the microscopic conductive and dielectric layers into circuits that define a chip’s final use and function. Its Customer Support Business Group (CSBG) provides products and services to maximize installed equipment performance and operational efficiency. Its customer base includes semiconductor memory, foundry, and integrated device manufacturers (IDMs) that make DRAM, NAND, and logic devices for these products.

Lam Research Careers

Joining Lam Research offers a unique opportunity to enhance your career at a leading global company at the forefront of innovation and technology in the semiconductor industry. Work You’ll Do At Lam Research, we are committed to advancing the technology landscape through continuous innovation and leadership. Our team of professionals is dedicated to pushing the boundaries of what's possible, making us a pivotal player in the semiconductor industry. By joining our team, you will collaborate with some of the brightest minds in the field, contributing to projects that significantly impact future technology. Transform Your Career Lam Research stands at the intersection of technology, innovation, and leadership. We offer job opportunities that challenge you to excel and push your limits. Our culture fosters growth and embraces diversity, ensuring that every team member can thrive professionally and personally. Join our dynamic team and be part of a company known for its groundbreaking work in developing equipment and services that drive the production of virtually every leading-edge chip in the world. Professional Growth and Development We believe in nurturing the professional growth of our employees by providing ample opportunities for career advancement through leadership and diversity training programs. Our commitment to your career is reflected in our robust benefits package, designed to support you and your family’s health, well-being, and financial future. Internship and Employment Opportunities Start your career path at Lam Research with our internship programs, which offer a hands-on experience in the semiconductor industry. Interns work closely with experienced mentors, gaining invaluable skills and knowledge that prepare them for full-time positions within our company. For seasoned professionals, we offer a variety of positions that leverage your skills to contribute to our mission of continuous improvement and innovation. We are always looking for curious, creative, and driven individuals to join our team. Inclusive Culture and Networking Lam Research is dedicated to creating a diverse and inclusive environment where all employees can thrive. Our culture encourages networking and collaboration, allowing you to connect with colleagues and industry leaders who are as passionate about technology and innovation as you are. Applying at Lam Research Ready to take the next step in your career? Explore the job opportunities at Lam Research by visiting our Careers page. Tailor your resume to highlight your relevant experience and skills, and prepare for an interview that could lead to a multitude of rewarding career paths with us. Stay Connected Keep up to date with the latest company news, employment trends, and career tips by joining our community. Subscribe to receive updates that can help you navigate your professional journey at Lam Research. Join Our Team Search open positions that match your skills and interests. We look for passionate, innovative, and solution-driven team players. Start your journey with Lam Research today, where your work isn’t just a job—it’s a pathway to personal and professional fulfillment. SEARCH LAM RESEARCH JOBS Discover the opportunities waiting for you at Lam Research, where we turn today’s innovations into tomorrow’s technologies.
Learn more about Lam Research
Size
14,100 employees
Market Cap
$54.8 billion
Industry
Net Income
$2.9 billion
Founded
2013
5 Year Trend
+16.5%
Revenue
$11.9 billion
NASDAQ

Similar Jobs

More Jobs at Lam Research

More Information Technology Jobs

Find similar Senior AI Observability engineer jobs: