Senior Site Reliability Engineer, Storage

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$168,000–$270,250 / yr
Posted
32 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $190k
This role $219k
$136k most similar roles pay here $285k

This role pays more than 76% of similar roles. Most pay $162,873–$217,725 — the shaded band above. At the midpoint, this role pays about $219k versus about $190k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Site Reliability Engineer, Storage

As a Senior Site Reliability Engineer - Storage, you will join the team focused on HPC storage to design, implement, and optimize on-prem High-Performance Computing infrastructure integrated with cloud computing. You will build distributed storage solutions, develop automation tools for deployment and monitoring, and ensure efficient operations of the growing IT ecosystem. Your daily work involves crafting scalable solutions for data-intensive applications, performing technology evaluations for distributed file systems, and collaborating with engineering teams to align infrastructure with developer needs. The role requires expertise in NetApp, Pure Storage, and S3-based storage like Cloudian MinIO, alongside experience with Lustre or GPFS filesystems. You will utilize Python, Bash, and Golang while operating within AWS, Azure, or GCP environments using monitoring stacks like Prometheus, Grafana, and Splunk to resolve performance bottlenecks for complex HPC applications.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Design and implement on-prem HPC storage solutions integrated with cloud computing to support data-intensive applications.
  • Develop automation tools for the deployment, management, monitoring, and alerting of large-scale infrastructure environments.
  • Manage enterprise NAS solutions including NetApp, Pure Storage, and S3-based storage like Cloudian or MinIO.
  • Administer and optimize distributed file systems such as Lustre or GPFS.
  • Perform technology evaluations and document best practices for distributed file systems and general procedures.
  • Identify and resolve performance bottlenecks for high-performance computing applications.
  • Influence methodologies for building, testing, and deploying applications to ensure optimal resource utilization.

What we're looking for

  • BS in Computer Science (or equivalent experience), MS with 5+ years of experience, or Ph.D. with 3 years of experience.
  • 8+ years of experience crafting technology solutions and resolving performance bottlenecks for HPC applications.
  • Experience designing, deploying, and managing Enterprise NAS solutions like NetApp, Pure Storage, and S3-based storage such as Cloudian MinIO.
  • Experience with one or more parallel or distributed filesystems such as Lustre or GPFS.
  • Proficiency in Python, Bash, or Golang programming and scripting.
  • Strong experience operating services in leading cloud environments including AWS, Azure, or GCP.
  • Experience with multiple monitoring stacks such as Prometheus+Grafana, Elasticsearch+Kibana, Splunk, or Zabbix.
  • Background with RDMA fabrics, HPC cluster management tools (Slurm, PBS, LSF), and containerization technologies like Docker or Kubernetes (preferred).

More like this

Similar roles

Senior Site Reliability Engineer, Storage

Nvidia

Santa Clara, CA 10 days ago $168,000$270,250
HPC Distributed File Systems Lustre GPFS NetApp Pure Storage S3 MinIO Python Bash Golang AWS Azure GCP Prometheus Grafana Elasticsearch Kibana Splunk Zabbix RDMA InfiniBand RoCE Slurm PBS LSF Docker Kubernetes
8+ yrs exp

Senior Manager, Storage Production Engineering

Nvidia

Remote 44 days ago $272,000$431,250
Lustre GPFS Ceph MinIO NetApp Pure Storage NVMe-oF RDMA NFS SMB iSCSI Fibre Channel Terraform Ansible Puppet Prometheus InfluxDB Elastic Stack Kubernetes S3 AWS Azure
10+ yrs exp Remote

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

Senior Manager, Storage Engineering

Nvidia

Santa Clara, CA 15 days ago $248,000$396,750
NVMe-oF NFS SMB/CIFS S3 GPFS Lustre NetApp Pure Storage Cloudian DDN Prometheus Grafana Ansible Python Bash REST APIs Kubernetes RDMA DPUs Slurm IBM LSF
10+ yrs exp

Senior Storage Production Engineer

Nvidia

Remote 37 days ago $176,000$276,000
Distributed Storage Kubernetes Python Go C/C++ Java Bash Terraform Ansible Chef Puppet Prometheus Grafana Elastic Stack InfluxDB CI/CD NFS SMB iSCSI S3 Fibre Channel RDMA NVMe over Fabrics Linux Git OpenStack
8+ yrs exp Remote

Senior AI Compute Engineer

Nvidia

Remote 151 days ago $148,000$235,750
Linux Python Bash Ansible Kubernetes SLURM LSF UGE InfiniBand MPI HPL NCCL MLPerf Lustre GPFS GPU Networking HPC System Administration Automation
8+ yrs exp Remote