Senior Staff Site Reliability Engineer, Compute Core Engineering

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CA
Salary
$200,000–$322,000 / yr
Posted
14 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $182k
This role $261k
$125k most similar roles pay here $343k

This role pays more than 93% of similar roles. Most pay $146,525–$217,556 — the shaded band above. At the midpoint, this role pays about $261k versus about $182k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Staff Site Reliability Engineer, Compute Core Engineering

As a Senior Staff Site Reliability Engineer - Compute Core Engineering, you will join the IT Compute Core Team to lead initiatives transforming architecture and building new service offerings across on-premise and cloud environments. You will design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP while ensuring high availability, capacity planning, and lifecycle management. Your daily work involves optimizing performance through software and hardware improvements like SR-IOV and DPU, as well as utilizing eBPF and XDP for observability and DDoS mitigation. You will develop tools for data visualization and reporting using Go or Python, manage infrastructure with Terraform and configuration management tools, and navigate complex network protocols including VLAN, VxLAN, SDN, BGP, and Anycast. This role focuses on building scalable distributed systems and containerized architectures to improve reliability across large bare-metal environments.

What does a Site Reliability Engineer earn in California?

Median $214000 from 54 postings across 15 companies.

See salary data

What you'll do

  • Lead initiatives to transform the IT Compute Core Team architecture and build new service offerings across on-prem and cloud environments.
  • Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP at a global scale.
  • Implement software and hardware optimizations such as SR-IOV and DPU to improve system performance and efficiency.
  • Utilize eBPF and XDP technologies for enhanced observability and DDoS mitigation.
  • Analyze system data to develop enterprise-wide capacity planning and infrastructure scaling strategies.
  • Develop and maintain custom tools for data collection, visualization, reporting, and automated alerting.
  • Evaluate existing application architectures to identify and implement containerization opportunities for improved scalability.
  • Manage large-scale bare metal environments using Infrastructure as Code (IaC) and configuration management tools.

What we're looking for

  • Bachelor's degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
  • 12+ years of proven experience in compute platform engineering with a focus on automation.
  • Experience designing and deploying containerization architectures and distributed systems infrastructure.
  • Proficiency in programming languages such as Go and/or Python.
  • Linux OS proficiency including knowledge of Kernel Internals.
  • Experience with Infrastructure as Code (IaC) tools like Terraform and configuration management tools.
  • Understanding of network protocols and architectures, including VLAN, VxLAN, SDN, BGP, and Anycast.
  • Experience managing large-scale environments consisting of BareMetal build infrastructure.

More like this

Similar roles

Senior Manager, Compute Core Engineering

Nvidia

Santa Clara, CA 12 days ago $248,000$391,000
Linux Go Python Terraform GitOps Distributed Systems DNS DHCP NTP PTP LDAP Container Platforms Infrastructure Automation SLIs SLOs HPC AI-assisted Operations
10+ yrs exp

Senior Staff Platform Engineer

Nvidia

Santa Clara, CA 23 days ago $200,000$322,000
Python Go Kubernetes Terraform AWS Azure Google Cloud Platform Infrastructure-as-Code Distributed Systems Linux TCP/IP DNS TLS HTTP/S CDN Load Balancing Content Delivery Capacity Analytics high-performance computing
10+ yrs exp

Senior Staff Site Reliability Engineer

Cisco

Remote (Irvine, CA) +1 15 days ago $192,400$275,800
SRE Linux Administration AWS GCP Azure Python Go Distributed Systems Splunk SPL Indexer Clustering Search Head Clusters KVStore Monitoring Alerting Observability Root Cause Analysis
10+ yrs exp Remote

Senior Site Reliability Engineer

Autodesk

Remote (ID) +1 28 days ago $117,000$209,330
Site Reliability Engineering AWS Kubernetes Python Go Java Infrastructure as Code CI/CD CloudWatch Splunk Datadog Dynatrace Bash PowerShell FedRAMP Distributed Systems Load Balancing DNS
7+ yrs exp Remote

Senior Site Reliability Engineer

Autodesk

San Francisco, CA 28 days ago $117,000$209,330
SRE Python Go Java Bash PowerShell AWS Kubernetes Infrastructure as Code CI/CD CloudWatch Splunk Datadog Dynatrace FedRAMP Distributed Systems Load Balancing DNS
7+ yrs exp