Principal Supercomputing Operations Software Engineer

Microsoft

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Salary
$142,800–$274,800 / yr
Posted
10 days ago
Freshness
Confirmed live yesterday
Closes
Feb 28, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $203k
This role $209k
$127k most similar roles pay here $291k

This role pays more than 61% of similar roles. Most pay $174,600–$232,187 — the shaded band above. At the midpoint, this role pays about $209k versus about $203k for comparable roles.

Based on 240 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 598 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 586 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Principal Supercomputing Operations Software Engineer

As a Principal Supercomputing Operations Software Engineer within the Artificial Intelligence and High Performance Computing organization, you will serve as the technical authority for InfiniBand and GPU interconnect fabric operations across large-scale AI supercomputing environments. You will manage high-severity incidents, perform multi-layer systems debugging across Subnet Manager, PCIe, firmware, drivers, and OS layers, and develop systemic prevention mechanisms to ensure training stability and SLA compliance. The role involves architecting automation, telemetry, and diagnostic tools to improve the observability of interconnect fabrics while authoring authoritative playbooks and escalation models. You will utilize skills in C, C++, C#, Java, JavaScript, or Python to solve complex reliability issues within high-performance computing infrastructure. Your work directly addresses the challenge of maintaining reliable, performant interconnect fabrics for frontier AI training and large-scale distributed simulations at scale.

What you'll do

  • Serve as the technical authority for InfiniBand and GPU interconnect fabric operations across large-scale AI supercomputing environments.
  • Lead and orchestrate high-severity fabric incidents including detection, triage, mitigation, recovery, and root cause analysis.
  • Perform multi-layer systems debugging across InfiniBand, Subnet Manager, PCIe, firmware, drivers, and OS layers.
  • Identify recurring failure patterns to define reliability models and systemic prevention mechanisms at fleet scale.
  • Author authoritative technical support guides, playbooks, and escalation frameworks for use across the organization.
  • Architect and develop automation, telemetry, and diagnostic tools to improve observability and mean time to mitigation.
  • Influence engineering direction by setting operational standards and partnering with hardware and firmware teams.

What we're looking for

  • Bachelor's Degree in Computer Science or related technical field and 6+ years of experience coding in languages like C, C++, C#, Java, JavaScript, or Python (or equivalent).
  • Ability to pass the Microsoft Cloud Background Check.
  • 6+ years of experience operating large-scale distributed systems, high-performance computing (HPC), or artificial intelligence (AI) infrastructure in production environments (preferred).
  • Demonstrated ownership of mission-critical production infrastructure with direct impact on service availability, GPU workloads, and customer SLAs (preferred).
  • Hands-on experience operating and debugging interconnect fabrics supporting large-scale compute workloads (preferred).
  • Strong Linux systems knowledge with experience debugging low-level infrastructure issues across operating systems, drivers, and services (preferred).
  • Proven ability to reason across hardware, firmware, drivers, and software stacks to diagnose and resolve complex production issues (preferred).
  • Bachelor's Degree in Computer Science or related field and 10+ years of engineering experience, OR a Master's Degree and 8+ years of experience (preferred).

More like this

Similar roles

Software Engineer, Compute Infra / HPC

Microsoft

37 days ago $142,800$274,800
Kubernetes Go Rust C++ C# Python Terraform Infrastructure as Code Distributed Systems Linux Azure InfiniBand RoCE RDMA NVLink NCCL Bare-metal Provisioning Container Runtimes
4+ yrs exp Hybrid

HPC Operations Engineer

Nvidia

Santa Clara, CA +3 4 days ago $124,000$195,500
Linux RHEL CentOS Ubuntu Python Bash Slurm LSF NFS LDAP HPC EDA Automation Scripting
2+ yrs exp Hybrid

Capacity & Efficiency Infrastructure

Microsoft

175 days ago $119,800$234,700
Python C++ CUDA Triton NCCL PyTorch JAX InfiniBand NVLink Distributed Training High-Performance Computing GPU Architecture LLM GenAI Telemetry Systems Profiling Benchmarking C# Java
6+ yrs exp Hybrid

Software Engineer, High Performance Computing

SpaceX

Hawthorne, CA 67 days ago $125,000$150,000
C++ Python C Linux ARM PowerPC x86 TCP UDP Machine Learning Unit Testing Hardware-in-the-loop Computer Architecture Networking Protocols Performance Optimization
2+ yrs exp

Senior Solutions Architect, Supercomputing

Nvidia

Remote (TX) +1 157 days ago $184,000$287,500
GPU CUDA Machine Learning Deep Learning High-Performance Computing Generative AI Agentic AI Docker Kubernetes Slurm GPGPU Parallel Computing Data Science Networking
8+ yrs exp Remote