Lead Principal Incident Manager, Data Centers

Oracle

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Nashville, TN
Salary
$110,200–$234,600 / yr
Posted
16 days ago
Freshness
Confirmed live today

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $185k
This role $172k
$95k most similar roles pay here $250k

This role pays less than 59% of similar roles. Most pay $154,437–$215,175 — the shaded band above. At the midpoint, this role pays about $172k versus about $185k for comparable roles.

Based on 240 similar postings.

Employer

About Oracle

Oracle Corporation is a leading multinational technology company specializing in database software, cloud computing, and enterprise software.

Oracle currently has 568 open roles on FindRole.

Listed pay typically runs $103,000–$234,600 across 425 roles with salary data.

Most-posted roles

View all roles at Oracle

At a glance

TL;DR · Lead Principal Incident Manager, Data Centers

As a Lead Principal Incident Manager, you will serve as a senior individual contributor leading the response to critical service-impacting events across Oracle Cloud Infrastructure data center infrastructure. You will coordinate site operations, infrastructure engineering, network engineering, platform reliability, security, and vendor teams to restore services quickly while managing incident timelines, decision logs, and stakeholder communications. The role involves performing root cause analysis, identifying systemic risks, and improving response playbooks and automation. You will utilize tools such as PagerDuty, ServiceNow, Jira, Datadog, Prometheus, and Grafana. Key competencies include incident command, operational decision-making, and cross-team coordination in high-pressure environments. The work focuses on maintaining the reliability of mission-critical cloud infrastructure, including compute, storage, networking, and facility-dependent services within a large-scale distributed environment involving high-density compute and high-performance networks.

What you'll do

  • Lead the response to critical incidents affecting OCI data center infrastructure including compute, storage, network, and facility services.
  • Create and maintain incident timelines, decision logs, action plans, and stakeholder communications during active events.
  • Provide concise, accurate updates to technical teams, operational leadership, and executive stakeholders during recovery efforts.
  • Coordinate with cross-functional engineering teams and external vendors to align responses across multiple infrastructure layers.
  • Lead post-incident reviews and root cause analysis to ensure corrective actions are tracked to completion.
  • Analyze incident data and recurring failure patterns to identify systemic risks and prioritize improvements.
  • Develop and maintain operational runbooks, incident command standards, and communication templates for repeatable response.
  • Track and report key performance metrics such as time to detect, mitigate, and resolve incidents.

What we're looking for

  • 6 to 10+ years of experience in incident management, infrastructure reliability engineering, data center operations, or other uptime-critical environments.
  • Experience leading high-severity incidents in large-scale distributed infrastructure, cloud, data center, or hybrid environments.
  • Strong understanding of data center operations, power, cooling, networking, and storage infrastructure.
  • Ability to lead cross-functional teams through high-pressure events while maintaining clear accountability.
  • Proficiency with incident management frameworks such as ITIL, SRE practices, or the Incident Command System.
  • Bachelor's degree in a technical field (preferred).
  • Experience supporting AI, HPC, or large-scale GPU infrastructure (preferred).
  • Formal training or certification in incident command, ITIL, service management, or related disciplines (preferred).

More like this

Similar roles

Senior Director Data Center Readiness

Oracle

Nashville, TN 107 days ago $169,800–$355,400
Oracle Cloud Infrastructure Data Center Infrastructure Safety Governance Site Readiness Infrastructure Reliability Continuous Improvement Compliance
10+ yrs exp

Director, Data Center Building Automation

Oracle

Nashville, TN 49 days ago
BMS SCADA PLC DCS HMI OT DCIM Schneider Electric Siemens Rockwell Ignition Niagara Johnson Controls Honeywell Cisco IEC 62443 NIST 800-82 Digital Twin EPMS
10+ yrs exp

Senior Manager, Incident Management

Zillow

Remote 59 days ago $132,400–$211,600
Incident Management Problem Management Root Cause Analysis (RCA) AI Workflows LLM SRE Technical Operations MTTR MTTD Runbooks JIRA ServiceNow Rootly Distributed Systems Cloud Infrastructure Automation
8+ yrs exp Remote

Principal Core Infrastructure Engineer

Oracle

Nashville, TN 24 days ago $114,600–$234,600
Oracle Cloud Infrastructure (OCI) GPU Grafana Unix Linux High Availability Incident Management root-cause analysis Automation Monitoring Alerting AI Agents
3+ yrs exp

Lead Principal Network Engineer

Oracle

Nashville, TN 51 days ago $146,300–$306,400
BGP OSPF IPv4 IPv6 VxLAN EVPN ECMP Python Git CI/CD Ansible Terraform YANG OpenConfig NETCONF SNMP gNMI NetFlow IPFIX TCP/IP
10+ yrs exp

Assistant Vice President, Incident Manager

Deutsche Bank

Jacksonville, FL 7 days ago $78,000–$120,500
ITIL Incident Management Knowledge Management AI Tools Service Improvement Governance Frameworks Technical Communication
Hybrid