Senior AI Observability Engineer

Lam Research

Confirmed live today High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Fremont, CA
Salary
$92,000–$211,000 / yr
Posted
30 days ago
Freshness
Confirmed live today
Closes
Feb 23, 2027

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $199k
This role $152k
$74k most similar roles pay here $261k

This role pays less than 82% of similar roles. Most pay $162,000–$235,750 — the shaded band above. At the midpoint, this role pays about $152k versus about $199k for comparable roles.

Based on 240 similar postings.

Employer

About Lam Research

Lam Research Corporation is a leading American supplier of wafer-fabrication equipment and services to the global semiconductor industry.

Lam Research currently has 314 open roles on FindRole.

Listed pay typically runs $114,000–$242,000 across 158 roles with salary data.

Most-posted roles

View all roles at Lam Research

At a glance

TL;DR · Senior AI Observability Engineer

As a Senior AI Observability engineer, you will join the IT team to build an AI-native operations layer for a hybrid enterprise estate spanning public and private clouds. You will serve as an individual contributor responsible for designing and shipping LLM-based agents, retrieval-augmented knowledge pipelines, and machine-learning anomaly detection systems to automate fault detection, triage, and remediation. Your daily work involves developing AIOps intelligence layers, engineering retrieval fabrics using vector stores, and managing LLMOps and MLOps workflows including prompt versioning and inference logging. You will utilize tools such as LangGraph, Semantic Kernel, AutoGen, and OpenTelemetry while applying expertise in SRE, network engineering, and cloud-agnostic design. The role solves complex infrastructure reliability problems by integrating automated incident response into the ITSM toolchain across diverse environments including manufacturing and high-performance computing facilities.

What you'll do

  • Build agentic AI workflows using LLM agents and orchestration frameworks for autonomous fault detection and remediation.
  • Develop an AIOps intelligence layer featuring time-series anomaly detection, alert deduplication, and predictive failure forecasting.
  • Engineer a retrieval knowledge fabric by indexing runbooks and technical documents into vector stores for RAG systems.
  • Implement AI-assisted incident response tools for automated summarization, root-cause analysis, and draft post-mortem generation.
  • Automate infrastructure remediation through event-driven pipelines with human-in-the-loop approval gates and audit trails.
  • Manage AI safety and governance including guardrails, hallucination monitoring, PII redaction, and prompt-injection defense.
  • Execute LLMOps and MLOps tasks such as model versioning, A/B testing, and tracking inference costs and latency.
  • Instrument AI systems with OpenTelemetry GenAI tracing and create dashboards for model performance and cost metrics.

What we're looking for

  • BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Eight or more years of experience in SRE, DevOps, infrastructure, observability, network engineering, or platform engineering.
  • Production experience with LLM applications including prompt engineering, RAG, embeddings, vector databases, and agent orchestration.
  • Experience using machine learning for anomaly detection, forecasting, event correlation, and alert noise reduction on operational telemetry.
  • Proficiency in at least one major AI platform with the ability to manage private-network deployment, quota, and cost.
  • Hands-on experience with at least one major public cloud and the ability to design portable, cloud-agnostic patterns.
  • Solid on-premises infrastructure background across virtualization, storage, data-center operations, and private cloud platforms.
  • Strong networking fundamentals including TCP/IP, BGP, OSPF, VLANs, MPLS, SD-WAN, firewalls, load balancers, and DNS.

More like this

Similar roles

GenAI Software Development Engineer

Amd

Santa Clara, CA +3 29 days ago $204,000–$306,000
LLM RAG Python TypeScript Go Rust Neo4j LightRAG React Next.js Terraform CI/CD AWS Azure GCP Microservices Distributed Systems LangGraph CrewAI AutoGen Model Context Protocol (MCP)

Senior ML/AI Engineer, Observability

General Motors (GM)

Sunnyvale, CA 89 days ago $178,420–$230,500
Go Kubernetes Docker Terraform Prometheus Grafana Istio GCP AWS Azure Unix Linux SSH Observability SLIs SLOs TSDBs
5+ yrs exp Hybrid

Senior Engineer, Observability Platform

CVS Health

Remote (RI) +4 41 days ago $83,430–$222,480
OpenTelemetry Go Python Java Spring Boot Kubernetes Docker Terraform CloudFormation Helm Kustomize Grafana Loki Tempo Mimir PostgreSQL MySQL Kafka Istio Envoy CI/CD
5+ yrs exp Remote

Senior Data Engineer, Observability Engineering

CVS Health

Remote (AZ) 6 days ago $92,700–$222,480
Databricks PySpark Spark SQL Python Delta Lake Unity Catalog Structured Streaming CI/CD GitHub Actions Azure DevOps OCSF Telemetry Distributed Data Processing Data Governance Metadata Management
5+ yrs exp Remote

Senior Software Engineer, Observability

Apple Inc

Cary, NC 142 days ago
OpenTelemetry Grafana Datadog Kotlin Go Python Java Kubernetes Prometheus CI/CD Terraform Pulumi LLMs NoSQL Distributed Systems SRE API Design Infrastructure-as-Code
7+ yrs exp