Senior Storage Production Engineer

Nvidia

Confirmed live 3 days ago High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$176,000–$276,000 / yr
Posted
90 days ago
Freshness
Confirmed live 3 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $188k
This role $226k
$126k most similar roles pay here $292k

This role pays more than 75% of similar roles. Most pay $149,762–$226,000 — the shaded band above. At the midpoint, this role pays about $226k versus about $188k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Storage Production Engineer

As a Senior Storage Production Engineer - DGX Cloud, you will join the team to design, implement, and support large-scale storage clusters ensuring high availability, data integrity, and low-latency access for AI/ML and HPC workloads. You will manage distributed systems by optimizing data placement, implementing intelligent caching, and automating operations through predictive analytics and infrastructure as code. Your daily work involves monitoring system health, managing capacity, and performing root cause analysis to ensure reliable GPU cloud services. The role requires expertise in Linux-based storage, parallel file systems, and protocols like NFS, S3, and NVMe over Fabrics. You will utilize tools including Ansible, Terraform, Prometheus, and Grafana while coding in languages such as Python, Go, C++, or Java to build automated frameworks for managing complex, high-performance distributed storage architectures at scale.

What you'll do

  • Design and support large-scale storage clusters to ensure scalability, high availability, and data integrity.
  • Develop monitoring, logging, and alerting systems for proactive detection of performance issues.
  • Optimize storage architectures specifically for AI/ML workloads to achieve low-latency and high-throughput performance.
  • Automate storage operations and infrastructure deployments using tools like Ansible, Terraform, or Python.
  • Manage capacity planning, data tiering strategies, and intelligent workload placement to improve efficiency.
  • Perform root cause analysis and participate in an on-call rotation to maintain system health.
  • Implement security measures including encryption and access controls for all storage systems.

What we're looking for

  • Hold a BS degree or equivalent experience in Computer Science, Storage Systems, or a related technical field.
  • Possess 8+ years of practical experience in the field.
  • Experience with distributed and high-performance storage solutions, including clustered/parallel file systems and distributed object storage.
  • Proficiency in storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
  • Expertise in algorithms, data structures, software design, and automating large-scale Linux-based storage systems.
  • Proficiency in at least one of the following languages: C/C++, Java, Python, Go, NodeJS, or Bash.
  • Experience with infrastructure configuration management tools like Ansible, Chef, Puppet, or Terraform.
  • Experience with observability and tracing tools such as InfluxDB, Prometheus, Grafana, and the Elastic stack.

More like this

Similar roles

Senior Storage Production Engineer

Nvidia

Remote 38 days ago $176,000$276,000
Distributed Storage Kubernetes Python Go C/C++ Java Bash Terraform Ansible Chef Puppet Prometheus Grafana Elastic Stack InfluxDB CI/CD NFS SMB iSCSI S3 Fibre Channel RDMA NVMe over Fabrics Linux Git OpenStack
8+ yrs exp Remote

Senior Manager, Storage Production Engineering

Nvidia

Remote 45 days ago $272,000$431,250
Lustre GPFS Ceph MinIO NetApp Pure Storage NVMe-oF RDMA NFS SMB iSCSI Fibre Channel Terraform Ansible Puppet Prometheus InfluxDB Elastic Stack Kubernetes S3 AWS Azure
10+ yrs exp Remote

Senior Storage Software Engineer

Nvidia

Remote 3 days ago $152,000$241,500
Go Python Rust C/C++ Java Kubernetes Linux eBPF POSIX FUSE Distributed Systems Cloud Infrastructure Telemetry Data Pipelines High-Performance Computing SRE
5+ yrs exp Remote

Senior Manager, Storage Engineering

Nvidia

Santa Clara, CA 16 days ago $248,000$396,750
NVMe-oF NFS SMB/CIFS S3 GPFS Lustre NetApp Pure Storage Cloudian DDN Prometheus Grafana Ansible Python Bash REST APIs Kubernetes RDMA DPUs Slurm IBM LSF
10+ yrs exp

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 33 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 11 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp