Senior Technical Program Manager, HPC

Microsoft

Confirmed live today Low trust

Quick summary

Work type
On-site
Location
—
Salary
$142,800–$274,800 / yr
Posted
2 days ago
Freshness
Confirmed live today
Closes
Mar 20, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $199k
This role $209k
$127k most similar roles pay here $291k

This role pays more than 63% of similar roles. Most pay $174,950–$222,226 — the shaded band above. At the midpoint, this role pays about $209k versus about $199k for comparable roles.

Based on 240 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 644 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 624 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Senior Technical Program Manager, HPC

Senior Technical Program Manager, HPC joins the Microsoft AI team to drive the operational health and reliability of high-performance computing infrastructure. This role focuses on ensuring production clusters remain available and ready to support demanding AI workloads by establishing operating mechanisms for cluster health, tracking service-level objectives, and managing critical incidents. The individual will coordinate cross-functional efforts across networking, storage, capacity, datacenter, and Azure teams to resolve dependencies and implement systemic fixes based on incident trends. Key responsibilities include defining operational metrics, managing production risks, and facilitating root-cause analysis. The role requires technical fluency in HPC, compute, accelerators, schedulers, and distributed systems. Candidates must possess experience in infrastructure engineering or site reliability to manage the complex interplay between hardware components and software layers within a large-scale high-performance computing environment.

What does a Technical Program Manager earn?

Median $191568 from 159 postings across 35 companies.

See salary data

What you'll do

  • Manage end-to-end programs for HPC production health, availability, and reliability.
  • Establish operating mechanisms to track cluster health, service level objectives, and infrastructure risks.
  • Lead coordinated responses for critical incidents including triage, escalation, and root-cause analysis.
  • Identify recurring reliability patterns and translate them into durable engineering improvements.
  • Build cross-functional partnerships across networking, storage, and datacenter teams to resolve dependencies.
  • Develop and track operational metrics and dashboards to provide visibility into cluster performance.
  • Proactively identify and mitigate production risks before they become critical blockers.
  • Communicate production health, major incidents, and risk trade-offs to leadership.

What we're looking for

  • Significant experience in technical program management, infrastructure engineering, production operations, site reliability, or large-scale distributed systems.
  • Demonstrated experience delivering complex, cross-functional programs spanning multiple engineering organizations.
  • Technical fluency in large-scale infrastructure including HPC, compute, accelerators, networking, storage, schedulers, capacity, telemetry, or distributed systems.
  • Experience driving production incident management, including escalation, mitigation, root-cause analysis, and corrective actions.
  • Experience working with SLAs, SLOs, availability metrics, and production-health indicators.
  • Ability to use operational data and incident trends to identify systemic problems and drive reliability improvements.
  • Strong ability to build relationships and influence engineering teams across organizational boundaries.
  • Excellent written and verbal communication skills to translate complex technical issues into clear risks and actions.
  • Experience supporting GPU or accelerator-based HPC/AI infrastructure at scale (preferred).
  • Experience with cloud infrastructure spanning compute, networking, storage, capacity, or datacenter operations (preferred).
  • Experience establishing reliability programs, operational scorecards, health dashboards, or incident-management mechanisms for large infrastructure environments (preferred).

More like this

Similar roles

Senior HW Technical Program Manager

Microsoft

14 days ago $119,800–$234,700
Technical Program Management Data Center Infrastructure Cloud Infrastructure Risk Management Product Development Data Analysis Debugging Supplier Management ODMs SIs IHVs
4+ yrs exp Hybrid

Senior Technical Program Manager

Microsoft

1 day ago $119,800–$234,700
Azure AI Data Analysis Technical Program Management Automation Infrastructure Security
4+ yrs exp Hybrid

Senior Technical Program Manager

Microsoft

1 day ago $119,800–$234,700
Quantum Computing Azure Agile Cryoelectronics High-Performance Computing Artificial Intelligence Firmware Control Systems ASICs Data Analysis Copilots Generative AI Risk Management
4+ yrs exp Hybrid

Senior HW Technical Program Manager

Microsoft

14 days ago $119,800–$234,700
Technical Program Management Data Center Infrastructure Risk Management Product Development Data Analysis Debugging Supplier Management ODMs SIs IHVs
4+ yrs exp Hybrid

Principal Supercomputing Operations Software Engineer

Microsoft

23 days ago $142,800–$274,800
InfiniBand GPU Interconnect high‑performance computing (HPC Linux Python C++ C# Java JavaScript PCIe Firmware Telemetry Root Cause Analysis Distributed Systems Automation
6+ yrs exp

Technical Program Manager II

Microsoft

8 days ago $102,100–$202,200
Azure Cloud Infrastructure Data Analysis Distributed Systems Datacenter Infrastructure Technical Program Management Product Development Risk Management Scalability Infrastructure Management
2+ yrs exp Hybrid