Principal Software Engineer, AI Infrastructure

Microsoft

Confirmed live today High trust
Hybrid

Quick summary

Work type
Hybrid
Location
—
Salary
$142,800–$274,800 / yr
Posted
5 days ago
Freshness
Confirmed live today
Closes
Apr 3, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $208k
This role $209k
$127k $291k
below market most similar roles pay here above market

This role pays more than 57% of similar roles. Most pay $174,600–$241,550 — the blue band above. At the midpoint, this role pays about $209k versus about $208k for comparable roles.

Based on 240 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 634 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 612 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Principal Software Engineer, AI Infrastructure

The Principal Software Engineer - AI Infrastructure will join the team to advance the reliability and observability architecture of a large-scale, globally distributed AI platform. This role involves defining the technical roadmap for platform health, designing unified telemetry capabilities, and developing safeguards for overload protection, capacity management, and routing integrity. The engineer will establish practices for production validation, fault testing, and automated rollbacks while leading cross-team architecture efforts and mentoring other engineers. Required skills include proficiency in coding languages such as C, C++, C#, Java, JavaScript, or Python. The work focuses on solving complex problems in distributed systems, cloud infrastructure, and AI inference, specifically addressing high-performance computing, accelerator-based infrastructure, and request tracing across distributed services to ensure operational health and reliability.

What does a Software Engineer earn?

Median $196750 from 2546 postings across 156 companies.

See salary data

What you'll do

  • Define and drive the architecture and technical roadmap for platform reliability, observability, and operational health.
  • Design unified health and telemetry capabilities connecting customer impact with infrastructure and deployment signals.
  • Develop platform safeguards for overload protection, capacity management, routing integrity, and automated recovery.
  • Establish engineering practices for production validation, progressive delivery, fault testing, and automated rollback.
  • Advance end-to-end request tracing and diagnostics across distributed services and asynchronous operations.
  • Lead cross-team architecture efforts and mentor engineers to turn production learnings into reusable platform capabilities.

What we're looking for

  • Bachelor's Degree in Computer Science or related technical field and 6+ years of technical engineering experience with coding in languages like C, C++, C#, Java, JavaScript, or Python.
  • Equivalent experience to the Bachelor's Degree and 6+ years of technical engineering experience is accepted.
  • Master's Degree in Computer Science or related technical field and 8+ years of technical engineering experience with coding in languages like C, C++, C#, Java, JavaScript, or Python (preferred).
  • Bachelor's Degree in Computer Science or related technical field and 12+ years of technical engineering experience with coding in languages like C, C++, C#, Java, JavaScript, or Python (preferred).
  • Experience with AI/ML serving platforms, high-performance computing, accelerator-based infrastructure, or other compute-intensive distributed systems (preferred).
  • Experience with distributed tracing, capacity management, load balancing, admission control, retries, backpressure, and graceful degradation (preferred).
  • Experience designing, building, and operating large-scale distributed systems, cloud services, or other complex production platforms (preferred).
  • Experience with reliability engineering, observability, service health, telemetry, and production incident response (preferred).

More like this

Similar roles

Principal Software Engineer, AI Frameworks

Microsoft

Remote 12 days ago $165,600–$296,400
C++ Python LLM Inference AI Frameworks Software Architecture Computer Architecture Kernel Optimization Azure Observability Benchmarking Compilers Runtimes AMD GPUs Microsoft Silicon
8+ yrs exp Remote

Senior Software Engineer, Responsible AI

Microsoft

61 days ago $119,800–$234,700
Azure AI Python C++ C# Java JavaScript C Azure OpenAI Azure ML Azure AI Studio CI/CD DevOps Distributed Systems Machine Learning SDLC Observability Telemetry
4+ yrs exp Hybrid

Principal Software Engineer, AI Compute Infrastructure

Arm Holdings

Seattle, WA 31 days ago $262,700–$355,400
Kubernetes Go Python Linux AWS EKS Terraform Argo CD Helm Prometheus Grafana PyTorch Ray vLLM SGLang TensorRT-LLM CUDA NVLink NVSwitch NCCL EFA DCGM Containers HPC
8+ yrs exp Hybrid

Principal Software Engineer, Core AI Platform

JPMorgan Chase

Seattle, WA 39 days ago
Python AI Infrastructure Machine Learning Distributed Systems Cloud Platforms System Design Observability APIs SDKs SRE Automated Testing Logging Metrics Incident Response Root-Cause Analysis Generative AI Infrastructure Engineering
10+ yrs exp

Principal AI Software Engineer

Amd

San Jose, CA 171 days ago $240,000–$360,000
ROCm CUDA OpenCL GPU Architecture CPU Architecture Hardware/Software Co-design Compilers Kernels Runtime AI/ML Frameworks Performance Optimization LLMs MoE Device Driver Development Operating Systems Open Source