Principal Software Developer, AI Infrastructure

Oracle

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Austin, TX
Salary
$114,600–$234,600 / yr
Posted
72 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $209k
This role $175k
$95k most similar roles pay here $294k

This role pays less than 75% of similar roles. Most pay $176,587–$241,550 — the shaded band above. At the midpoint, this role pays about $175k versus about $209k for comparable roles.

Based on 240 similar postings.

Employer

About Oracle

Oracle Corporation is a leading multinational technology company specializing in database software, cloud computing, and enterprise software.

Oracle currently has 781 open roles on FindRole.

Listed pay typically runs $102,300–$209,500 across 713 roles with salary data.

Most-posted roles

View all roles at Oracle

At a glance

TL;DR · Principal Software Developer, AI Infrastructure

Principal Software Developer, AI Infrastructure joins the OCI AI Infrastructure team to develop a high-performance GPU platform supporting AI, ML, and HPC workloads. This role involves designing and implementing software and firmware for managing GPU-based servers, including architectural changes for delivery, health monitoring, triage automation, and diagnostic services. The candidate will build systems that allow users to scale across thousands of GPUs using RoCE and Infiniband technologies. Key responsibilities include troubleshooting, debugging, and collaborating with partner teams to manage and repair GPU systems. Required skills include proficiency in languages like Java, Python, C, C++, or Go, along with expertise in Linux systems, distributed systems, and algorithms. The role addresses the technical challenge of managing large-scale production environments, ensuring fault tolerance, state management, and high-performance data planes for massive computing clusters.

What does a Software Developer earn?

Median $151475 from 90 postings across 16 companies.

See salary data

What you'll do

  • Design, implement, and deliver software and firmware for managing GPU-based AI servers.
  • Develop architectural changes for GPU delivery, health monitoring, triage automation, and diagnostic services.
  • Build high-performance systems to support distributed AI/ML/HPC workloads across thousands of GPUs.
  • Troubleshoot and debug complex issues within the low-level system stack and infrastructure.
  • Manage large-scale production systems involving over 1000 server instances.
  • Develop solutions utilizing RoCE and Infiniband networking technologies for high-performance computing.
  • Resolve customer issues by collaborating with product teams to identify and fix technical bugs.

What we're looking for

  • BS or MS degree in Computer Science or a related technical field involving coding.
  • 6+ years of experience delivering and operating large-scale production systems with over 1,000 server instances.
  • Deep understanding of operating systems, computer networks, and high-performance applications.
  • Proficiency in at least one programming language such as Java, Python, C, C++, Go, or Shell scripting.
  • Proven ability to deliver products through the full software development lifecycle.
  • Experience with Linux systems and system-level architecture, including fault tolerance and state management.
  • Familiarity with GPU hardware architecture, Infiniband, RoCE networking, and public cloud service data planes.
  • Knowledge of databases (MySQL), caching technologies (Redis, Memcache), and distributed systems algorithms.

More like this

Similar roles

Principal Software Engineer, AI Infra Compute

Oracle

Austin, TX +1 84 days ago $114,600$234,600
Python Java TypeScript OCI AWS Azure GCP Kafka Docker Linux Bash Perl Ruby RESTful APIs Swagger OpenAPI NLP RoCE Infiniband Agile
6+ yrs exp

Software Developer V, AI Infrastructure

Oracle

Seattle, WA 80 days ago $135,200$306,400
Python Java C C++ Shell Scripting Linux OCI RoCE Infiniband GPU Distributed Systems Networking Firmware SmartNICs Agile Bare Metal
10+ yrs exp

Senior Software Developer, AI Infra Compute

Oracle

Austin, TX +1 73 days ago $89,200$209,500
Oracle Cloud Infrastructure GPU RoCE Infiniband Java Python C C++ Shell Scripting Linux Distributed Systems Firmware Agile Networking
4+ yrs exp

Software Developer IV

Oracle

Seattle, WA 98 days ago $114,600$234,600
OCI Python Java C C++ Shell Scripting Linux Infiniband RoCE SQL MySQL Redis Memcache Distributed Systems Firmware Agile
6+ yrs exp

Software Engineer, AI Infrastructure

Apple Inc

Cupertino, CA 56 days ago $184,700$324,800
Python Go LLMs Distributed Systems Cloud Services Security Architecture AI Infrastructure Observability Load Shaping Cost Controls Backend Programming
5+ yrs exp