Senior HPC Storage Engineer

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CAAustin, TX
Salary
$184,000–$287,500 / yr
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $191k
This role $236k
$129k most similar roles pay here $304k

This role pays more than 85% of similar roles. Most pay $155,900–$226,000 — the shaded band above. At the midpoint, this role pays about $236k versus about $191k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior HPC Storage Engineer

As a Senior HPC Storage Engineer on the HW Infrastructure Storage Strategy team, you will provide leadership in researching, designing, and implementing groundbreaking fast storage solutions for demanding high-performance computing and computationally intensive workloads. You will identify architectural changes for file, block, and object storage to meet scaling requirements of expanding cloud infrastructure while performing capacity modeling and growth planning. Your daily responsibilities include developing automation tools for large-scale management, conducting technology evaluations for distributed file systems, and performing root cause analysis. The role requires expertise in Linux distributions like CentOS or Ubuntu, Python programming, bash scripting, and container technologies such as Docker and Enroot. You will address technical challenges involving parallel filesystems like Ceph, Lustre, or GPFS, while optimizing deep learning workflows using PyTorch and TensorFlow across high-performance networking environments.

What you'll do

  • Research and design scalable, next-generation distributed storage services for high-performance computing workloads.
  • Optimize storage performance and cost-effectiveness to meet growing infrastructure needs for large-scale environments.
  • Develop automation tools for managing large-scale infrastructure, including monitoring, alerting, and self-service resource consumption.
  • Perform technology evaluations and define procedures related to distributed file systems.
  • Analyze and optimize deep learning workflows and perform performance analysis for researchers.
  • Conduct root cause analysis and implement corrective actions for storage issues at various scales.
  • Guide methodologies for building, testing, and deploying applications to ensure efficient resource utilization.

What we're looking for

  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 8+ years of experience designing and/or operating large scale storage infrastructure.
  • Experience analyzing and tuning storage performance for various workloads.
  • Proficiency in CentOS, RHEL, and/or Ubuntu Linux distributions.
  • Proficiency in Python programming and bash scripting.
  • In-depth understanding of container technologies such as Docker and Enroot.
  • Expertise in parallel and distributed filesystems like Ceph, Weka.io, Vast, Lustre, or GPFS.
  • Familiarity with NVIDIA GPUs, CUDA programming, NCCL, and deep learning frameworks like PyTorch and TensorFlow.

More like this

Similar roles

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 10 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 32 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior Manager, Storage Production Engineering

Nvidia

Remote 44 days ago $272,000$431,250
Lustre GPFS Ceph MinIO NetApp Pure Storage NVMe-oF RDMA NFS SMB iSCSI Fibre Channel Terraform Ansible Puppet Prometheus InfluxDB Elastic Stack Kubernetes S3 AWS Azure
10+ yrs exp Remote

Senior Manager, Storage Engineering

Nvidia

Santa Clara, CA 15 days ago $248,000$396,750
NVMe-oF NFS SMB/CIFS S3 GPFS Lustre NetApp Pure Storage Cloudian DDN Prometheus Grafana Ansible Python Bash REST APIs Kubernetes RDMA DPUs Slurm IBM LSF
10+ yrs exp

Senior Storage Production Engineer

Nvidia

Remote 89 days ago $176,000$276,000
Distributed Storage Kubernetes Python Go C/C++ Java Bash Ansible Terraform Prometheus Grafana CI/CD NFS S3 iSCSI NVMe over Fabrics Linux Infrastructure as Code
8+ yrs exp Remote

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote