Senior AI Performance Architect

Microsoft

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
—
Salary
$119,800–$234,700 / yr
Posted
4 days ago
Freshness
Confirmed live yesterday
Closes
Apr 4, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $199k
This role $177k
$105k $257k
below market most similar roles pay here above market

This role pays less than 65% of similar roles. Most pay $162,000–$235,750 — the blue band above. At the midpoint, this role pays about $177k versus about $199k for comparable roles.

Based on 240 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 634 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 612 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Senior AI Performance Architect

The Senior AI Performance Architect joins the Systems Planning and Architecture team within the Azure Hardware Systems and Infrastructure organization. In this role, you will analyze frontier training and inference workloads to identify performance bottlenecks and develop analytical models, simulators, and data-analysis tools. You will evaluate trade-offs in scalability, utilization, and cost across accelerator, memory, networking, and software configurations. Key responsibilities include running targeted microbenchmarks on silicon to measure compute and communication performance, and guiding hardware-software co-design decisions for cloud-scale AI systems. Required skills include experience in computer architecture, performance modeling, and systems analysis. Preferred technical proficiencies include Python, C++, CUDA, Triton, or ROCm/HIP, alongside knowledge of transformer-based models, HBM hierarchies, and distributed execution strategies like tensor or pipeline parallelism.

What you'll do

  • Analyze frontier training and inference workloads to identify performance requirements and bottlenecks.
  • Develop and improve analytical models, simulators, and data-analysis tools for AI system performance.
  • Evaluate trade-offs in performance, scalability, utilization, capacity, and cost across hardware and software configurations.
  • Translate workload insights into architecture requirements by studying hardware-software interactions in large-scale systems.
  • Guide hardware-software co-design decisions for cloud-scale AI systems with MSI and partner teams.
  • Run targeted microbenchmarks on silicon to measure compute, memory, communication, and synchronization performance.
  • Compare analytical model predictions with silicon measurements to identify and close performance gaps.
  • Communicate findings and recommendations through data visualizations, technical reports, and architecture decision materials.

What we're looking for

  • Master's Degree in a related field and 3+ years of technical engineering experience, or Bachelor's Degree and 5+ years, or equivalent experience.
  • Ability to pass the Microsoft Cloud background check.
  • Doctorate in a related field and 3+ years of experience, or Master's and 6+ years, or Bachelor's and 8+ years (preferred).
  • Experience in AI systems performance, computer architecture, performance modeling, high-performance computing, distributed systems, or accelerator software (preferred).
  • Strong understanding of computer architecture and performance interactions among compute, memory, communication, and software (preferred).
  • Experience with analytical performance modeling, simulation, profiling, workload characterization, or system performance analysis (preferred).
  • Hands-on programming experience in Python and at least one systems or accelerator programming environment such as C++, CUDA, Triton, or ROCm/HIP (preferred).
  • Experience with transformer-based models, large language model training/inference, or distributed execution strategies (preferred).

More like this

Similar roles

Senior AI Training Performance Architect

Nvidia

Santa Clara, CA 79 days ago $184,000–$287,500
CUDA C++ Python GPU Architecture Deep Learning Neural Networks Performance Modeling MLPerf Training Computer Architecture System Simulators Performance Analysis Drivers DL Frameworks
5+ yrs exp

Senior Solutions Architect, AI Performance Engineering

Nvidia

Santa Clara, CA 86 days ago $184,000–$287,500
CUDA C++ Python Linux GPU Kernel Optimization cuDNN cuBLAS CUTLASS NCCL NVSHMEM InfiniBand RoCE NVLINK DeepGEMM FlashAttention Nsight Compute High-Performance Computing Parallel Programming
8+ yrs exp

Senior AI Systems Performance Engineer

Amd

Austin, TX 4 days ago $178,400–$267,600
PyTorch ONNX Runtime vLLM TensorFlow ROCm Ryzen AI Python Shell Linux BIOS Profiling Tracing Performance Engineering Computer Vision Generative AI

Senior AI/ML Performance Engineer

General Motors (GM)

Sunnyvale, CA +1 5 days ago $144,700–$261,300
Python PyTorch Kubernetes AWS GCP Azure Nvidia DCGM nvidia-smi Grafana BigQuery Hugging Face Nvidia Nsight Nsight Compute HPC Distributed Systems GPU Architecture
5+ yrs exp Hybrid

Senior AI-Native Architect

Carnegie Mellon University

Pittsburgh, PA 4 days ago
Software Architecture AI Agentic Workflows Reference Architectures Technical Strategy Prototyping
10+ yrs exp

Senior Accelerated Computing Architect

Nvidia

Santa Clara, CA 180 days ago $184,000–$287,500
CUDA C++ C OpenCL MPI NVSHMEM OpenSHMEM Python IPC Linear Algebra Numerical Methods Software Design Algorithms GPU Programming Multi-node Communication
6+ yrs exp