Senior Manager, Storage Production Engineering

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$272,000–$431,250 / yr
Posted
44 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $191k
This role $352k
$112k most similar roles pay here $465k

This role pays more than 99% of similar roles. Most pay $155,450–$226,175 — the shaded band above. At the midpoint, this role pays about $352k versus about $191k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Manager, Storage Production Engineering

As the Senior Manager, Storage Production Engineering, you will lead and coach a team of engineers to ensure storage systems remain fast, reliable, and scalable for advanced AI and high-performance computing workloads. You will manage day-to-day operations including capacity planning, data lifecycle management, incident response, and root cause analysis while collaborating with engineering and AI teams to optimize data pipelines. The role involves designing and improving distributed storage, parallel file systems, and object storage using automation, monitoring, and analytics. Key technical requirements include experience with Lustre or GPFS, Ceph or MinIO, and enterprise platforms like NetApp or Pure Storage. You will work with technologies such as NVMe over Fabrics, RDMA, Terraform, Ansible, Prometheus, and the Elastic stack to manage infrastructure for complex data movement and high-availability storage environments.

What you'll do

  • Lead and coach a team of Storage Production Engineers in a collaborative and learning-focused environment.
  • Design, deploy, and improve large-scale storage systems including distributed storage, parallel file systems, and object storage.
  • Utilize automation, monitoring, and analytics to improve the reliability and efficiency of storage services.
  • Manage capacity planning, data lifecycle management, cost awareness, and disaster recovery plans for storage infrastructure.
  • Evaluate and adopt modern technologies such as NVMe over Fabrics, RDMA, and high-speed interconnects.
  • Lead incident response and root cause analysis to implement permanent fixes for recurring storage issues.
  • Partner with engineering and AI/ML teams to optimize data pipelines and workflow performance.

What we're looking for

  • BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience.
  • Experience in large scale storage architecture, operations, production engineering, or infrastructure.
  • 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.
  • Experience managing infrastructure operations including on-call rotations, incident response, troubleshooting, and managing SLOs/KPIs.
  • Hands-on experience with parallel file systems (Lustre, GPFS), distributed storage (Ceph, MinIO), and enterprise object or NAS platforms.
  • Strong knowledge of block, file, and object storage performance tuning, data protection, and high availability design.
  • Experience with storage networking and protocols including NFS, SMB, iSCSI, Fibre Channel, RDMA, and NVMe-oF.
  • Practical experience with automation/infrastructure as code (Terraform, Ansible, Puppet) and monitoring tools (Prometheus, InfluxDB, Elastic stack).

More like this

Similar roles

Senior Manager, Storage Engineering

Nvidia

Santa Clara, CA 15 days ago $248,000$396,750
NVMe-oF NFS SMB/CIFS S3 GPFS Lustre NetApp Pure Storage Cloudian DDN Prometheus Grafana Ansible Python Bash REST APIs Kubernetes RDMA DPUs Slurm IBM LSF
10+ yrs exp

Senior Storage Production Engineer

Nvidia

Remote 37 days ago $176,000$276,000
Distributed Storage Kubernetes Python Go C/C++ Java Bash Terraform Ansible Chef Puppet Prometheus Grafana Elastic Stack InfluxDB CI/CD NFS SMB iSCSI S3 Fibre Channel RDMA NVMe over Fabrics Linux Git OpenStack
8+ yrs exp Remote

Senior Storage Production Engineer

Nvidia

Remote 89 days ago $176,000$276,000
Distributed Storage Kubernetes Python Go C/C++ Java Bash Ansible Terraform Prometheus Grafana CI/CD NFS S3 iSCSI NVMe over Fabrics Linux Infrastructure as Code
8+ yrs exp Remote

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 10 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 32 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior HPC Storage Engineer

Nvidia

Santa Clara, CA +1 3 days ago $184,000$287,500
Distributed Storage HPC Python Bash Docker Enroot Ceph Weka.io Vast Lustre GPFS CUDA NCCL MLPerf NVMe SDN PyTorch TensorFlow Linux RHEL
8+ yrs exp