Engineering Manager, AI Platform & SRE

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$208,000–$333,500 / yr
Employment
Full-time
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $229k
This role $271k
$166k most similar roles pay here $351k

This role pays more than 75% of similar roles. Most pay $185,000–$273,149 — the shaded band above. At the midpoint, this role pays about $271k versus about $229k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1371 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 1096 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Engineering Manager, AI Platform & SRE

Engineering Manager – AI Platform & SRE As an Engineering Manager within the Site Reliability Engineering team, you will lead a group of engineers responsible for building and operating resilient AI platform capabilities at enterprise scale. You will define technical strategies and roadmaps to deliver highly available, scalable distributed systems that support enterprise AI agent products. Your daily work involves driving the development of AI agents and intelligent automation for incident response, troubleshooting, and remediation while improving developer experience through infrastructure-as-code and self-service platforms. The role requires expertise in Linux, networking, Kubernetes, and public cloud platforms like AWS, Azure, or GCP. You will utilize languages such as Python, Go, TypeScript, JavaScript, or Java to manage production services. Key technical focuses include observability via OpenTelemetry, capacity management, and establishing service-level objectives to ensure reliable infrastructure for complex AI operations.

What does a Engineering Manager earn in California?

Median $290250 from 62 postings across 17 companies.

See salary data

What you'll do

  • Lead and develop a team of SRE, platform, and software engineers for the AI Platform Runtime.
  • Define technical strategies, priorities, and roadmaps aligned with product and business objectives.
  • Oversee the design and delivery of highly available, scalable, and secure distributed systems for enterprise AI products.
  • Drive the development of AI agents and intelligent automation for incident response and platform operations.
  • Establish measurable reliability goals using SLIs, SLOs, error budgets, and capacity models.
  • Improve developer experience through self-service platforms, infrastructure-as-code, and standardized delivery patterns.
  • Manage critical incidents and conduct blameless postmortems to implement durable corrective actions.
  • Recruit, coach, and mentor engineers to foster a high-performing and inclusive team environment.

What we're looking for

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, Software Engineering, Cloud Infrastructure, or a related technical field.
  • 3+ years managing or formally leading engineering teams responsible for complex production systems.
  • Technical foundation in distributed systems, Linux, networking, Kubernetes, and public cloud platforms like AWS, Azure, or GCP.
  • Experience leading teams building production software and automation using Python, Go, TypeScript, JavaScript, or Java.
  • Solid understanding of observability at scale, including OpenTelemetry, metrics, logs, distributed tracing, profiling, and operational analytics.
  • Experience applying SRE practices such as SLOs, error budgets, capacity management, incident management, and blameless postmortems.
  • Demonstrated ability to create technical roadmaps, manage competing priorities, and deliver measurable outcomes across multiple teams.
  • Hands-on experience with AI agents, agentic workflows, or intelligent automation for infrastructure (preferred); experience scaling SRE functions in large enterprises (preferred).

More like this

Similar roles

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 26 days ago $248,000–$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK CloudFormation AI/ML Platforms High-Performance Computing
10+ yrs exp Hybrid

Staff Site Reliability Engineer, AI Platform Runtime

Nvidia

Santa Clara, CA 26 days ago $168,000–$270,250
Site Reliability Engineering Kubernetes Python Go TypeScript JavaScript Terraform AWS CDK CloudFormation CrossPlane OpenTelemetry infrastructure-as-code Observability Distributed Systems Networking AWS Azure GCP
10+ yrs exp Hybrid

Engineering Manager, AI Data Platforms & Quality

Apple Inc

Cupertino, CA 64 days ago $237,600–$356,400
Python SQL Spark Airflow Kafka RAG LLM Kubernetes Docker AWS GCP Azure CI/CD Distributed Systems Data Pipelines Vector Search Metadata Management Observability

Engineering Manager, Service Platform

Instacart

Remote (CA, Canada) +4 129 days ago $196,000–$207,000
Temporal Canary Deployment Tooling Distributed Systems Infrastructure Internal Tools Developer Experience
7+ yrs exp Remote

Engineering Manager, AI Developer Technology

Nvidia

Santa Clara, CA +4 60 days ago $224,000–$356,500
Deep Learning Machine Learning LLMs Computer Vision CUDA C++ C OpenMP MPI pthreads GPU Architecture Parallel Programming Linear Algebra Performance Optimization Multimodal Architectures
8+ yrs exp Hybrid