Lead Principal Site Reliability Engineer

Oracle

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Vienna, VA
Salary
$96,300–$264,100 / yr
Posted
53 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $185k
This role $180k
$76k most similar roles pay here $284k

This role pays less than 51% of similar roles. Most pay $156,312–$213,817 — the shaded band above. At the midpoint, this role pays about $180k versus about $185k for comparable roles.

Based on 240 similar postings.

Employer

About Oracle

Oracle Corporation is a leading multinational technology company specializing in database software, cloud computing, and enterprise software.

Oracle currently has 781 open roles on FindRole.

Listed pay typically runs $102,300–$209,500 across 713 roles with salary data.

Most-posted roles

View all roles at Oracle

At a glance

TL;DR · Lead Principal Site Reliability Engineer

Lead Principal Site Reliability Engineer serves as a technical leader and individual contributor focused on ensuring high availability and operational excellence for mission-critical applications within a cloud platform. The role involves designing and architecting infrastructure, managing capacity planning, overseeing incident response, and implementing automated solutions to reduce manual toil. You will build and maintain CI/CD pipelines, develop Infrastructure as Code using Terraform, and create automation scripts utilizing Python, Bash, or PowerShell. Key technical requirements include experience with Kubernetes, Docker, Linux administration, and Oracle Cloud Infrastructure. The position requires expertise in observability tools like Prometheus, Grafana, and Splunk to monitor metrics, logs, and traces. You will solve complex problems regarding scalability and performance while ensuring systems meet strict reliability standards within highly regulated environments, ultimately improving the resilience of critical business services through proactive engineering and robust infrastructure management.

What you'll do

  • Maintain high availability, reliability, and performance for enterprise applications and cloud infrastructure.
  • Design and implement automation to reduce manual operational effort and ensure deployment consistency.
  • Develop Infrastructure as Code (IaC) using Terraform and scripts in Python or Bash.
  • Build and maintain CI/CD pipelines to support automated deployments and release management.
  • Monitor systems using observability platforms including metrics, logs, traces, and alerting.
  • Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Lead incident response, conduct root cause analyses, and implement corrective actions.
  • Perform capacity planning, performance tuning, and scalability assessments for cloud-native applications.

What we're looking for

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • 8+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Engineering.
  • Experience supporting production environments with strict availability requirements.
  • Hands-on experience with Oracle Cloud Infrastructure (OCI) or other major cloud providers like AWS, Azure, or GCP.
  • Experience with Kubernetes, Docker, and container orchestration platforms.
  • Proficiency in Infrastructure as Code (Terraform preferred).
  • Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
  • Strong scripting skills in Python, Bash, or PowerShell and experience with Linux system administration.

More like this

Similar roles

Lead Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $96,300$264,100
Site Reliability Engineering Kubernetes Docker Terraform Ansible Chef Puppet Python Go Java JavaScript Bash Oracle Cloud Infrastructure Microsoft Azure Google Cloud Platform infrastructure-as-code Chaos Engineering
6+ yrs exp

Principal Site Reliability Engineer

Oracle

Reston, VA +1 45 days ago $84,900$209,500
Kubernetes Terraform Docker Python Bash Linux Unix Oracle Database RAC Chef Puppet DNS DHCP HTTP TCP/IP LLM VMware Cisco
6+ yrs exp

Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $84,900$209,500
Site Reliability Engineering Oracle Cloud Infrastructure Linux Windows Server Python PowerShell Bash Ansible Chef infrastructure-as-code Networking DNS Firewalls Load Balancing Certificates incident-management Observability Capacity Planning
3+ yrs exp

Site Reliability Engineer, Lead

Booz Allen Hamilton

Chantilly, VA 50 days ago $99,000$225,000
Prometheus Grafana ELK Stack Linux AWS Python Terraform Terragrunt Kubernetes OpenTelemetry AWS CloudWatch AWS EKS Rancher Jenkins Git Docker Nessus JIRA Confluence SRE
8+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Palo Alto, CA 59 days ago
Site Reliability Engineering Java Go Python Terraform Kubernetes Docker CI/CD GitOps Grafana Prometheus Dynatrace Datadog Splunk Kafka RabbitMQ SQS Neo4j Pinecone Weaviate Chroma LangChain LangGraph AutoGen CrewAI GitHub Copilot Fluentd Logstash Vector RESTful APIs RAG TensorFlow PyTorch scikit-learn Hadoop Spark Flink MongoDB Cassandra DynamoDB InfluxDB TimescaleDB AWS Azure GCP
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp