Principal Systems Engineer

Oracle

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Nashville, TN
Posted
10 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $186k
$131k most similar roles pay here $252k

This listing doesn't post a salary. Most similar roles pay $158,000–$214,500.

Based on 240 similar postings.

Employer

About Oracle

Oracle Corporation is a leading multinational technology company specializing in database software, cloud computing, and enterprise software.

Oracle currently has 571 open roles on FindRole.

Listed pay typically runs $100,150–$209,500 across 474 roles with salary data.

Most-posted roles

View all roles at Oracle

At a glance

TL;DR · Principal Systems Engineer

As a Principal Systems Engineer on the AI Infra Operations team, you will lead the development and maintenance of automation and operational tooling for GPU infrastructure within Oracle Cloud Infrastructure across multiple geographic regions. You will be responsible for building robust monitoring, alerting, and diagnostics using Grafana to ensure high availability, while serving as a senior escalation point for complex host issues and participating in incident response. The role involves managing large-scale distributed systems, performing hardware lifecycle tasks like provisioning and repair, and developing AI agents. Required technical skills include expert Linux administration (Ubuntu and Oracle Linux), Python, Bash, and experience with NVIDIA or AMD based systems. You will collaborate with engineering and hardware teams to improve GPU fleet health, capacity, and regional deployment readiness while mentoring junior engineers on operational best practices.

What does a Systems Engineer earn?

Median $152000 from 248 postings across 37 companies.

See salary data

What you'll do

  • Build and maintain automation and operational tooling for OCI GPU infrastructure across multiple geographic regions.
  • Develop monitoring, alerting, and diagnostic systems for GPU fleet health, performance, and capacity using tools like Grafana.
  • Serve as the senior escalation point for complex GPU host issues and hardware repairs.
  • Lead incident response and root-cause analysis to resolve blockers affecting GPU availability and regional deployments.
  • Improve AI2 operations processes, GPU fleet automation, and OCI region build readiness.
  • Develop and monitor the safe rollout of AI agents and associated tooling.
  • Document operational procedures, automation workflows, troubleshooting guides, and runbooks.
  • Mentor junior engineers on operational best practices and provide senior technical support.

What we're looking for

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 8+ years of software operations or infrastructure automation experience with proficiency in Python and Bash.
  • Expert Linux administration experience in large-scale production environments, specifically with Ubuntu and Oracle Linux.
  • Strong understanding of distributed systems including peer-to-peer, node-to-node, and service-to-service communication patterns.
  • Experience with data-center and host-lifecycle tasks such as provisioning, validation, repair workflows, and fleet recovery.
  • Experience with observability tooling for metrics, logging, dashboards, and alerting.
  • Experience with AI agents and tooling.
  • Experience leading on-call operations and incident response; experience with GPU infrastructure (preferred).

More like this

Similar roles

Senior Core Infrastructure Engineer

Oracle

Nashville, TN 19 days ago $79,200–$209,500
Oracle Cloud Infrastructure (OCI) Linux GPU Automation Incident Management Observability Public Cloud Scripting Networking Operating Systems
3+ yrs exp

Principal Software Engineer, DGX Cloud Production Engineering

Nvidia

Remote (Santa Clara, CA) 139 days ago $272,000–$431,250
Kubernetes Go Python GitOps Linux GPU Clusters AI/ML Infrastructure Distributed Systems Infrastructure Automation APIs Observability SLOs Multi-cloud BMaaS VMaaS High-Performance Computing
10+ yrs exp Remote

Principal Software Engineer

The Walt Disney Company

Remote 111 days ago $206,400–$276,700
AV1 HEVC VVC AVC Python C++ Deep Learning Computer Vision GStreamer Docker REST APIs Git GitHub XML JSON ABR Algorithms Video Processing
10+ yrs exp Remote

Principal Software Engineer

General Dynamics

San Antonio, TX 77 days ago $119,000–$161,000
DevSecOps CI/CD Kubernetes Docker Terraform Ansible Puppet Chef Python Java Go Bash AWS Azure Google Cloud Agile Methodology infrastructure-as-code OWASP Nessus Burp Suite Splunk
10+ yrs exp Hybrid

Principal Software Engineer

Microsoft

Redmond, WA 64 days ago $165,600–$296,400
LLM C# C++ Python Java JavaScript Distributed Systems Azure AI Microsoft Copilot Orchestration Runtime Architecture Platform Engineering MCP Declarative Programming Async Workflows Agentic Workflows
8+ yrs exp