Principal Engineer, Cloud Site Reliability Engineering

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$272,000–$431,250 / yr
Posted
37 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $193k
This role $352k
$94k most similar roles pay here $467k

This role pays more than 99% of similar roles. Most pay $172,000–$213,564 — the shaded band above. At the midpoint, this role pays about $352k versus about $193k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Engineer, Cloud Site Reliability Engineering

Principal Engineer, Cloud Site Reliability Engineering joins the Infrastructure, Planning and Process Cloud Infrastructure Team to serve as an SRE Architect for the GPU Private Cloud. This role involves architecting, implementing, and supporting end-to-end CI/CD systems using open-source and proprietary software while optimizing critical software development workflows across various internal organizations. The engineer will identify performance bottlenecks, improve cost efficiency for AI development and testing systems, and lead software projects while technically directing a team of engineers. Key responsibilities include onboarding internal teams to cloud infrastructure and crafting metrics via analytics dashboards. Required skills include Java, Python, Shell-script, REST APIs, and experience with SQL/NoSQL databases like MySQL, Cassandra, MongoDB, or Elasticsearch. Technical expertise must include Docker, Kubernetes, OpenStack, Chef, Puppet, Hadoop, Ceph, SwiftStack, LXC, Git, Perforce, JFrog, and Kafka to support diverse operating systems and hardware platforms.

What does a Engineer earn in California?

Median $208000 from 94 postings across 22 companies.

See salary data

What you'll do

  • Architect, implement, and support end-to-end CI/CD systems using open-source and proprietary software.
  • Develop software solutions to optimize critical development workflows across various internal organizations.
  • Identify performance bottlenecks to improve the speed and cost efficiency of AI development systems.
  • Lead software development projects and provide technical direction to a team of engineers.
  • Onboard internal development teams to private cloud infrastructure based on specific use cases.
  • Create and implement critical metrics using various analytics methods and dashboards.
  • Resolve complex issues within distributed software systems and large-scale infrastructure.

What we're looking for

  • BS or MS in Electrical Engineering, Computer Science, or a related field (or equivalent experience).
  • 15+ years of systems software development experience.
  • At least 1 year of experience dedicated to developing or exploring AI.
  • Experience maintaining cloud infrastructure and highly available production environments.
  • Proficiency in Java, Python, and Shell scripting with an understanding of distributed systems and REST APIs.
  • Experience with SQL/NoSQL database systems such as MySQL, Cassandra, MongoDB, or Elasticsearch.
  • Expertise in Docker containers, Virtual Machines, and technologies like OpenStack, Kubernetes, Chef/Puppet, and Kafka.
  • Ability to collaborate across organizational boundaries in a multi-national, multi-time-zone corporate environment.

More like this

Similar roles

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 3 days ago $248,000$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK CloudFormation AI/ML Platforms High-Performance Computing
10+ yrs exp Hybrid

Senior Lead Site Reliability Engineer

JPMorgan Chase

Palo Alto, CA 59 days ago
Site Reliability Engineering Java Go Python Terraform Kubernetes Docker CI/CD GitOps Grafana Prometheus Dynatrace Datadog Splunk Kafka RabbitMQ SQS Neo4j Pinecone Weaviate Chroma LangChain LangGraph AutoGen CrewAI GitHub Copilot Fluentd Logstash Vector RESTful APIs RAG TensorFlow PyTorch scikit-learn Hadoop Spark Flink MongoDB Cassandra DynamoDB InfluxDB TimescaleDB AWS Azure GCP
5+ yrs exp

Senior Solutions Architect, IPP

Nvidia

Remote 91 days ago $224,000$356,500
Python Java Kubernetes Docker CI/CD OpenStack MySQL Cassandra MongoDB Elasticsearch Kafka Hadoop Ceph SwiftStack Chef Puppet Git Perforce JFrog REST APIs Linux Windows Android
10+ yrs exp Remote

Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $84,900$209,500
Site Reliability Engineering Oracle Cloud Infrastructure Linux Windows Server Python PowerShell Bash Ansible Chef infrastructure-as-code Networking DNS Firewalls Load Balancing Certificates incident-management Observability Capacity Planning
3+ yrs exp

Principal Site Reliability Engineer

Oracle

Reston, VA +1 45 days ago $84,900$209,500
Kubernetes Terraform Docker Python Bash Linux Unix Oracle Database RAC Chef Puppet DNS DHCP HTTP TCP/IP LLM VMware Cisco
6+ yrs exp

Senior Engineer, Edge-Cloud Reliability

Anduril Industries

Reston, VA 72 days ago $191,000$253,000
Go Rust Python GCP NixOS Keycloak LDAP SRE platform-engineering distributed-systems DNS Networking Identity Federation RBAC Bare-metal