Senior Site Reliability Engineer, Storage

Nvidia

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$168,000–$270,250 / yr
Posted
10 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $190k
This role $219k
$136k most similar roles pay here $285k

This role pays more than 76% of similar roles. Most pay $162,873–$217,725 — the shaded band above. At the midpoint, this role pays about $219k versus about $190k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Site Reliability Engineer, Storage

As a Senior Site Reliability Engineer - Storage, you will join the team focused on HPC storage to design, implement, and optimize on-prem High-Performance Computing infrastructure integrated with cloud computing. You will be responsible for crafting distributed storage solutions, building automation tools for deployment and monitoring, and ensuring efficient operations of the growing IT ecosystem. Your daily work involves developing tooling for self-service resource consumption, performing technology evaluations for distributed file systems, and collaborating with engineering teams to align infrastructure with developer workflows. The role requires expertise in Enterprise NAS solutions like NetApp and Pure Storage, S3-based storage such as Cloudian MinIO, and parallel filesystems including Lustre or GPFS. You will utilize Python, Bash, and Golang while managing environments across AWS, Azure, or GCP using monitoring stacks like Prometheus, Grafana, Elasticsearch, Kibana, Splunk, and Zabbix.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Design and implement on-prem HPC storage solutions integrated with cloud computing to support growing IT needs.
  • Build scalable and efficient storage solutions tailored for data-intensive applications while optimizing performance and cost.
  • Develop automation tools for the deployment, management, monitoring, and alerting of large-scale infrastructure environments.
  • Create self-service mechanisms for resource consumption across the organization's infrastructure.
  • Perform technology evaluations and document procedures related to distributed file systems.
  • Identify and resolve performance bottlenecks for high-performance computing applications.
  • Manage enterprise NAS solutions including NetApp, Pure Storage, and S3-based storage like Cloudian MinIO.
  • Oversee parallel or distributed filesystems such as Lustre and GPFS.

What we're looking for

  • BS in Computer Science with 8+ years of experience, MS with 5+ years, or PhD with 3+ years of experience.
  • 8+ years of experience crafting technology solutions and resolving performance bottlenecks for HPC applications.
  • Experience designing, deploying, and managing Enterprise NAS solutions like NetApp, Pure Storage, and S3-based storage such as Cloudian MinIO.
  • Experience with parallel or distributed filesystems such as Lustre or GPFS.
  • Proficiency in Python, Bash, or Golang programming and scripting.
  • Strong experience operating services in major cloud environments including AWS, Azure, or GCP.
  • Experience with monitoring stacks such as Prometheus+Grafana, Elasticsearch+Kibana, Splunk, or Zabbix.
  • Preferred experience with RDMA fabrics, HPC cluster management tools (Slurm, PBS, LSF), and containerization technologies like Kubernetes.

More like this

Similar roles

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 32 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior Manager, Storage Production Engineering

Nvidia

Remote 44 days ago $272,000$431,250
Lustre GPFS Ceph MinIO NetApp Pure Storage NVMe-oF RDMA NFS SMB iSCSI Fibre Channel Terraform Ansible Puppet Prometheus InfluxDB Elastic Stack Kubernetes S3 AWS Azure
10+ yrs exp Remote

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

Senior Manager, Storage Engineering

Nvidia

Santa Clara, CA 15 days ago $248,000$396,750
NVMe-oF NFS SMB/CIFS S3 GPFS Lustre NetApp Pure Storage Cloudian DDN Prometheus Grafana Ansible Python Bash REST APIs Kubernetes RDMA DPUs Slurm IBM LSF
10+ yrs exp

Senior Storage Production Engineer

Nvidia

Remote 37 days ago $176,000$276,000
Distributed Storage Kubernetes Python Go C/C++ Java Bash Terraform Ansible Chef Puppet Prometheus Grafana Elastic Stack InfluxDB CI/CD NFS SMB iSCSI S3 Fibre Channel RDMA NVMe over Fabrics Linux Git OpenStack
8+ yrs exp Remote

Senior AI Compute Engineer

Nvidia

Remote 151 days ago $148,000$235,750
Linux Python Bash Ansible Kubernetes SLURM LSF UGE InfiniBand MPI HPL NCCL MLPerf Lustre GPFS GPU Networking HPC System Administration Automation
8+ yrs exp Remote