Lead Principal Site Reliability Engineer

Oracle

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Nashville, TN
Salary
$96,300–$264,100 / yr
Posted
57 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $187k
This role $180k
$76k most similar roles pay here $284k

This role pays less than 54% of similar roles. Most pay $160,259–$214,375 — the shaded band above. At the midpoint, this role pays about $180k versus about $187k for comparable roles.

Based on 238 similar postings.

Employer

About Oracle

Oracle Corporation is a leading multinational technology company specializing in database software, cloud computing, and enterprise software.

Oracle currently has 781 open roles on FindRole.

Listed pay typically runs $102,300–$209,500 across 713 roles with salary data.

Most-posted roles

View all roles at Oracle

At a glance

TL;DR · Lead Principal Site Reliability Engineer

Lead Principal Site Reliability Engineer serves as a technical authority and consultant, leading the design and architecture of infrastructure and services for large-scale, distributed, and business-critical platforms. The role involves owning capacity forecasting, collaborating with software development teams to build scalable infrastructures, and overseeing incident response and maintenance tasks. You will develop observability strategies including monitoring, logging, tracing, and alerting while implementing automation to reduce operational toil through infrastructure-as-code and automated remediation. Key technologies include major cloud platforms like Oracle Cloud Infrastructure, AWS, Azure, or Google Cloud Platform; containerization tools such as Docker and Kubernetes; and configuration management tools like Terraform, Ansible, Chef, or Puppet. You will utilize programming languages including Python, Go, Java, JavaScript, or Bash to solve complex problems regarding system reliability, performance, and high-availability for mission-critical systems.

What you'll do

  • Design and architect scalable infrastructure to ensure system reliability and performance standards.
  • Define and implement observability strategies including SLIs, SLOs, error budgets, and automated alerting.
  • Develop automation to reduce manual toil and improve deployment, configuration management, and recovery processes.
  • Lead technical response and coordination during high-severity production incidents and conduct post-incident reviews.
  • Perform capacity planning, demand forecasting, and load testing for critical platforms.
  • Conduct resilience testing through failure-mode analysis and controlled fault-injection exercises.
  • Partner with security teams to integrate compliance and vulnerability remediation into infrastructure operations.
  • Provide technical mentorship and leadership to engineers across the organization on reliability best practices.

What we're looking for

  • Candidates must have 6 to 10+ years of experience in site reliability, software engineering, or cloud infrastructure.
  • Experience is required in designing and operating highly available production systems at significant scale.
  • Deep knowledge of distributed systems, cloud architecture, networking, operating systems, storage, and databases is required.
  • Advanced experience with a major cloud platform like OCI, AWS, Azure, or GCP is required.
  • Proficiency with containerization and orchestration technologies including Docker and Kubernetes is required.
  • Proven expertise in infrastructure-as-code and configuration management tools such as Terraform or Ansible is required.
  • Strong programming or scripting skills in languages such as Python, Go, Java, JavaScript, or Bash are required.
  • A Bachelor’s or advanced degree in computer science, engineering, information systems, or a related field is preferred.

More like this

Similar roles

Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $84,900$209,500
Site Reliability Engineering Oracle Cloud Infrastructure Linux Windows Server Python PowerShell Bash Ansible Chef infrastructure-as-code Networking DNS Firewalls Load Balancing Certificates incident-management Observability Capacity Planning
3+ yrs exp

Principal Site Reliability Engineer

Oracle

Reston, VA +1 45 days ago $84,900$209,500
Kubernetes Terraform Docker Python Bash Linux Unix Oracle Database RAC Chef Puppet DNS DHCP HTTP TCP/IP LLM VMware Cisco
6+ yrs exp

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 3 days ago $248,000$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK CloudFormation AI/ML Platforms High-Performance Computing
10+ yrs exp Hybrid

Lead Site Reliability Engineer

Alloy

New York, NY 158 days ago $179,000$250,000
Kubernetes AWS Terraform Python Go Docker Infrastructure as Code Datadog CloudWatch ELK EFK Distributed Systems SRE
10+ yrs exp Hybrid

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 43 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 43 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp