Software Engineer, Compute Infra / HPC

Microsoft

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Salary
$142,800–$274,800 / yr
Posted
37 days ago
Freshness
Confirmed live yesterday
Closes
Feb 1, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $202k
This role $209k
$127k most similar roles pay here $291k

This role pays more than 64% of similar roles. Most pay $168,500–$235,750 — the shaded band above. At the midpoint, this role pays about $209k versus about $202k for comparable roles.

Based on 240 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 598 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 586 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Software Engineer, Compute Infra / HPC

Software Engineer - Compute Infra / HPC will join the Microsoft AI team to build compute infrastructure software that ensures frontier AI systems remain healthy at global scale. The role involves designing and building distributed services, control planes, and platform APIs for the end-to-end lifecycle of AI compute clusters, including provisioning, certification, and automated recovery from hardware or software failures. You will develop foundational primitives such as Kubernetes control planes, networking, and infrastructure-as-code to automate cluster bring-up and fleet health monitoring. Key technologies include Go, Rust, C++, C#, Java, or Python, alongside expertise in Kubernetes, Linux systems, and high-performance interconnects like InfiniBand or RoCE. This position addresses the challenge of transforming raw capacity into secure, observable, and schedulable clusters while replacing manual operational workflows with durable software capabilities to improve researcher velocity and fleet availability.

What does a Software Engineer earn?

Median $197500 from 2196 postings across 133 companies.

See salary data

What you'll do

  • Design and build distributed services, control planes, and platform APIs for the end-to-end lifecycle of AI compute clusters.
  • Develop foundational primitives for cluster bootstrap, Kubernetes control planes, networking, identity, and infrastructure as code.
  • Automate rack and node qualification, topology validation, and scale testing to reduce cluster bring-up time.
  • Build a fleet health platform to detect, isolate, and remediate hardware, network, and storage issues automatically.
  • Create policy-driven automation for the certification, maintenance, and safe rollout of software and infrastructure components.
  • Translate recurring operational incidents into durable software abstractions and automated guardrails to eliminate manual toil.
  • Use production evidence, benchmarks, and fault injection to improve system architecture and recovery times.
  • Provide technical leadership through design reviews, architectural decisions, and mentoring for complex cross-team initiatives.

What we're looking for

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
  • 4+ years of software engineering experience building production distributed systems, infrastructure platforms, or control-plane services.
  • 4+ years of experience designing and implementing scalable software for cloud, datacenter, cluster, or fleet infrastructure using Go, Rust, C++, C#, Java, or Python.
  • Experience with Kubernetes, Linux systems, public-cloud infrastructure, and infrastructure-as-code or declarative configuration.
  • Demonstrated ability to debug complex behavior across service, control-plane, operating-system, networking, and hardware boundaries.
  • Experience leading projects across multiple teams and communicating technical tradeoffs to various stakeholders.
  • Preferred: Master's degree in Computer Science or a related technical field and 6+ years of experience building distributed systems or infrastructure platforms.
  • Preferred: Experience with GPU/accelerator systems, high-performance interconnects (InfiniBand, RoCE, RDMA), and low-level Linux system components.

More like this

Similar roles

Software Engineer, Compute

Apple Inc

Seattle, WA 9 days ago $142,300$263,300
Site Reliability Engineering Go Python Kubernetes OpenStack KVM Terraform Ansible Chef Infrastructure as Code Unix Distributed Systems Bare Metal VM Orchestration Automation
1+ yrs exp

Software Engineer - Compute

Apple Inc

Seattle, WA 2 days ago $142,300$263,300
Site Reliability Engineering Go Python Kubernetes Terraform Ansible Chef OpenStack KVM Infrastructure as Code Distributed Systems Unix Bare Metal VM Orchestration Automation
1+ yrs exp

Software Engineer, High Performance Computing

SpaceX

Hawthorne, CA 67 days ago $125,000$150,000
C++ Python C Linux ARM PowerPC x86 TCP UDP Machine Learning Unit Testing Hardware-in-the-loop Computer Architecture Networking Protocols Performance Optimization
2+ yrs exp

System Software Engineer, HPC Performance

Nvidia

Champaign, IL +2 17 days ago $152,000$241,500
C C++ Python HPC Machine Learning Deep Learning Artificial Intelligence x86 ARM Linux Windows macOS Profiling Tools Cloud Computing
5+ yrs exp