Principal Site Reliability Engineer

Oracle

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Reston, VAAustin, TX
Salary
$84,900–$209,500 / yr
Posted
45 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $185k
This role $147k
$67k most similar roles pay here $247k

This role pays less than 79% of similar roles. Most pay $156,886–$214,000 — the shaded band above. At the midpoint, this role pays about $147k versus about $185k for comparable roles.

Based on 238 similar postings.

Employer

About Oracle

Oracle Corporation is a leading multinational technology company specializing in database software, cloud computing, and enterprise software.

Oracle currently has 781 open roles on FindRole.

Listed pay typically runs $102,300–$209,500 across 713 roles with salary data.

Most-posted roles

View all roles at Oracle

At a glance

TL;DR · Principal Site Reliability Engineer

As a Principal Site Reliability Engineer, you will join a dynamic team providing cloud operations for Oracle National Security Realms. You will solve complex infrastructure problems by designing, writing, and deploying software to improve the availability, scalability, and efficiency of products and services. Your daily responsibilities include managing high-impact incidents as an escalation point, mentoring junior engineers, automating manual tasks to reduce toil, and performing capacity planning and system tuning for large-scale distributed systems. You will work with Linux and Unix operating systems using Docker, Kubernetes, and Terraform. Required skills include proficiency in scripting languages like Python, Bash, or Perl, knowledge of core protocols such as DNS and TCP/IP, and experience with configuration management tools like Chef or Puppet. The role specifically involves managing AI infrastructure, including clustered GPUs and LLM deployment.

What does a Site Reliability Engineer earn?

Median $186200 from 131 postings across 36 companies.

See salary data

What you'll do

  • Solve complex infrastructure problems and build automation to prevent recurring issues in cloud services.
  • Design, develop, and deploy software to improve the availability, scalability, and efficiency of Oracle products.
  • Create and maintain production architectures, system designs, and standards for large-scale distributed systems.
  • Identify manual tasks and implement automation to reduce toil and enable continuous delivery.
  • Serve as an escalation point for junior engineers during high-impact incidents and provide technical mentorship.
  • Manage complex change management tickets while ensuring safety and minimal disruption to services.
  • Perform capacity planning, demand forecasting, and system tuning for cloud operations.
  • Deploy and maintain AI infrastructure including clustered GPUs and LLM integrations.

What we're looking for

  • Must have 6 to 10+ years of experience in site reliability engineering or related fields.
  • Must be a U.S. citizen possessing and maintaining a TS/SCI with Poly security clearance.
  • Must hold a technology-related bachelor's degree or equivalent work experience.
  • Proficiency in Linux and Unix operating systems, including deep knowledge of internals and host-based networking.
  • Proficiency in scripting languages such as Python, Bash, Ruby, Perl, JavaScript, or Java for task automation.
  • Experience with Docker, Kubernetes, Terraform, and configuration management tools like Chef or Puppet.
  • Familiarity with core protocols (DNS, DHCP, HTTP, TCP/IP) and experience managing monitoring solutions for large-scale environments.
  • Specific experience in deploying AI infrastructure, including clustered GPUs and LLM deployment and maintenance.

More like this

Similar roles

Lead Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $96,300$264,100
Site Reliability Engineering Kubernetes Docker Terraform Ansible Chef Puppet Python Go Java JavaScript Bash Oracle Cloud Infrastructure Microsoft Azure Google Cloud Platform infrastructure-as-code Chaos Engineering
6+ yrs exp

Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $84,900$209,500
Site Reliability Engineering Oracle Cloud Infrastructure Linux Windows Server Python PowerShell Bash Ansible Chef infrastructure-as-code Networking DNS Firewalls Load Balancing Certificates incident-management Observability Capacity Planning
3+ yrs exp

Senior Site Reliability Engineer

Oracle

Reston, VA +1 45 days ago $81,100$187,000
Kubernetes Terraform Docker Python Bash Linux Unix Chef Puppet DNS DHCP HTTP TCP/IP LLM Deployment Cloud Computing Configuration Management
3+ yrs exp

Senior Site Reliability Engineer

Oracle

Reston, VA +1 50 days ago $81,100$187,000
Kubernetes Docker Terraform Python Go Java Bash Shell Ruby Perl JavaScript Linux Unix Chef Puppet DNS DHCP HTTP TCP LLM Deployment
3+ yrs exp

Senior Site Reliability Engineer

Oracle

Reston, VA +1 50 days ago $81,100$187,000
Linux Unix Docker Kubernetes Terraform Python Go Shell Bash Ruby Perl Java JavaScript Chef Puppet DNS DHCP HTTP TCP GPU LLM
3+ yrs exp

Senior Site Reliability Engineer

Oracle

Reston, VA +1 50 days ago $81,100$187,000
Docker Kubernetes Terraform Python Go Shell Java Ruby Perl JavaScript Linux Unix Chef Puppet DNS DHCP HTTP TCP GPU LLM
3+ yrs exp