Principal Site Reliability Engineer, Compute Infrastructure

Palo Alto Networks

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$175,400–$283,800 / yr
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $192k
This role $230k
$122k most similar roles pay here $301k

This role pays more than 84% of similar roles. Most pay $168,725–$215,687 — the shaded band above. At the midpoint, this role pays about $230k versus about $192k for comparable roles.

Based on 240 similar postings.

Employer

About Palo Alto Networks

Palo Alto Networks is a global cybersecurity company offering network security, cloud security, and endpoint protection solutions.

Palo Alto Networks currently has 40 open roles on FindRole.

Listed pay typically runs $164,000–$266,000 across 38 roles with salary data.

Most-posted roles

View all roles at Palo Alto Networks

At a glance

TL;DR · Principal Site Reliability Engineer, Compute Infrastructure

Principal Site Reliability Engineer, Compute Infrastructure joins the Information Technology team as an AI-Native Reliability Software Engineer. This role focuses on building software engines, intelligent pipelines, and autonomous systems to power a cloud presence by treating the cloud as a programmable, AI-orchestrated entity. The engineer will lead enterprise architectural strategy across AWS, Azure, and GCP while integrating AI workflows and LLM agents for automated infrastructure refactoring. Key responsibilities include architecting secure-by-design telemetry pipelines, engineering autonomous agents for predictive auto-scaling, and managing globally distributed Kubernetes fleets with AI-driven traffic routing. The role requires expertise in Go, Python, or TypeScript (Node.js), along with experience in vector databases, prompt engineering, and CI/CD tools like GitHub Actions or GitLab CI to solve complex problems regarding high-availability cloud infrastructure and automated incident response.

What does a Site Reliability Engineer earn in California?

Median $214000 from 59 postings across 16 companies.

See salary data

What you'll do

  • Lead enterprise architectural strategy across AWS, Azure, and GCP while integrating AI workflows and LLM agents.
  • Architect secure-by-design telemetry pipelines using LLMs for dynamic IAM auditing and automated vulnerability patching.
  • Engineer autonomous agents in Go, Python, or TypeScript to create predictive auto-scaling and self-healing systems.
  • Design internal AI agents to perform autonomous root-cause analysis and resolve complex system anomalies.
  • Define the technical vision for globally distributed Kubernetes fleets with AI-driven traffic routing and capacity planning.
  • Build intelligent delivery pipelines featuring automated testing, security gates, and AI-assisted code reviews.
  • Partner with Core AI/ML teams to bridge the gap between model deployment and high-availability cloud infrastructure.

What we're looking for

  • 8+ years of experience in Cloud Software Engineering, SRE, or Distributed Systems Infrastructure (BS or equivalent).
  • Proven track record with large-scale distributed systems, multi-cloud platforms (GCP, AWS), and container orchestration.
  • Strong software engineering fundamentals in TypeScript (Node.js), Go, or Python.
  • At least 2+ years of hands-on experience integrating AI tools, LLMs, or predictive analytics into deployment workflows (preferred).
  • Experience interfacing with LLM APIs, vector databases, and prompt engineering for systems-level orchestration (preferred).
  • Experience designing agentic SRE workflows for autonomous incident response, root-cause analysis, and self-healing of distributed infrastructure (preferred).
  • Experience building intelligent delivery pipelines using GitHub Actions or GitLab CI with automated testing and security gates (preferred).
  • Ability to partner with Core AI/ML teams to bridge the gap between model deployment and high-availability cloud infrastructure (preferred).

More like this

Similar roles

Staff Software Engineer, Cloudops

Palo Alto Networks

Santa Clara, CA 31 days ago $124,500–$201,300
AWS Azure GCP Kubernetes Terraform Pulumi AWS CDK Python Go TypeScript Node.js LangChain CrewAI OpenTelemetry Prometheus GitHub Actions GitLab CI LLMs Vector Databases Linux CI/CD
6+ yrs exp

Senior Staff Software Engineer, AI Platform

Palo Alto Networks

Santa Clara, CA 17 days ago $153,700–$248,600
LLM Python TypeScript Go Node.js OpenAI Anthropic Google Vertex AI AWS Bedrock AWS Azure GCP Kubernetes Terraform Pulumi AWS CDK CI/CD GitHub Actions GitLab CI Distributed Systems
8+ yrs exp

Principal Engineer, Cloud Site Reliability Engineering

Nvidia

Santa Clara, CA 59 days ago $272,000–$431,250
SRE CI/CD Python Java Kubernetes Docker OpenStack MySQL Cassandra MongoDB Elasticsearch Kafka Hadoop Ceph SwiftStack Chef Puppet Git Perforce JFrog REST APIs Linux Windows Android
10+ yrs exp

Principal Site Reliability Engineer

Oracle

Nashville, TN 78 days ago $84,900–$209,500
Site Reliability Engineering Oracle Cloud Infrastructure Linux Windows Server Python PowerShell Bash Ansible Chef infrastructure-as-code Networking DNS Firewalls Load Balancing Certificates incident-management Observability Capacity Planning
3+ yrs exp

Principal Site Reliability Engineer

Oracle

Nashville, TN 78 days ago $84,900–$209,500
Site Reliability Engineering Linux Windows Server Python Bash PowerShell Ansible Chef Oracle Cloud Infrastructure infrastructure-as-code Monitoring Logging Observability Capacity Planning incident-management Patching Citrix
3+ yrs exp