Senior Storage Production Engineer

Nvidia

Confirmed live 2 days ago High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$176,000–$276,000 / yr
Posted
37 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $188k
This role $226k
$126k most similar roles pay here $292k

This role pays more than 75% of similar roles. Most pay $149,762–$226,000 — the shaded band above. At the midpoint, this role pays about $226k versus about $188k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Storage Production Engineer

As a Senior Storage Production Engineer - DGX Cloud, you will join the team to design, implement, and support large-scale storage clusters ensuring high availability, data integrity, and low-latency access for AI/ML workloads. You will manage distributed systems by developing monitoring, logging, and alerting tools while optimizing performance through compression, deduplication, and intelligent caching. The role involves automating storage operations using Python, Go, C/C++, Java, or Bash, alongside infrastructure tools like Ansible, Chef, Puppet, and Terraform. You will work with technologies including Kubernetes, NVMe over Fabrics, RDMA, S3, and the Elastic stack to manage high-throughput systems. Your daily work focuses on solving complex data management challenges, ensuring reliable GPU cloud services, and maintaining production infrastructure through proactive fault detection, capacity planning, and automated remediation within a large-scale distributed storage environment.

What you'll do

  • Design and support large-scale storage clusters to ensure high availability, scalability, and data integrity.
  • Develop monitoring, logging, and alerting systems to proactively detect and resolve performance issues.
  • Optimize storage architectures for AI/ML workloads to achieve low-latency access and high-throughput performance.
  • Automate storage operations and infrastructure deployments using tools like Ansible, Terraform, or Python.
  • Manage capacity planning, data tiering strategies, and intelligent workload placement to improve efficiency.
  • Maintain production infrastructure by monitoring system health, latency, and availability through predictive analytics.
  • Implement security measures including encryption, access controls, and auditing for all storage systems.
  • Participate in an on-call rotation to provide incident response and perform root cause analysis.

What we're looking for

  • BS degree or equivalent experience in Computer Science, Storage Systems, or a related technical field.
  • 8+ years of practical experience in the field.
  • Experience with distributed and high-performance storage solutions including clustered/parallel file systems and object storage.
  • Proficiency in storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
  • Expertise in algorithms, data structures, software design, and automating large-scale Linux-based storage systems.
  • Experience with programming languages including C/C++, Java, Python, Go, NodeJS, or Bash for automation and performance tuning.
  • Hands-on experience with infrastructure configuration management tools like Ansible, Chef, Puppet, or Terraform.
  • Experience with observability and tracing tools such as InfluxDB, Prometheus, Grafana, and the Elastic stack.

More like this

Similar roles

Senior Storage Production Engineer

Nvidia

Remote 89 days ago $176,000$276,000
Distributed Storage Kubernetes Python Go C/C++ Java Bash Ansible Terraform Prometheus Grafana CI/CD NFS S3 iSCSI NVMe over Fabrics Linux Infrastructure as Code
8+ yrs exp Remote

Senior Manager, Storage Production Engineering

Nvidia

Remote 44 days ago $272,000$431,250
Lustre GPFS Ceph MinIO NetApp Pure Storage NVMe-oF RDMA NFS SMB iSCSI Fibre Channel Terraform Ansible Puppet Prometheus InfluxDB Elastic Stack Kubernetes S3 AWS Azure
10+ yrs exp Remote

Senior Storage Software Engineer

Nvidia

Remote 2 days ago $152,000$241,500
Go Python Rust C/C++ Java Kubernetes Linux eBPF POSIX FUSE Distributed Systems Cloud Infrastructure Telemetry Data Pipelines High-Performance Computing SRE
5+ yrs exp Remote

Senior Manager, Storage Engineering

Nvidia

Santa Clara, CA 15 days ago $248,000$396,750
NVMe-oF NFS SMB/CIFS S3 GPFS Lustre NetApp Pure Storage Cloudian DDN Prometheus Grafana Ansible Python Bash REST APIs Kubernetes RDMA DPUs Slurm IBM LSF
10+ yrs exp

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 32 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 10 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp