Principal AI Software Engineer

Microsoft

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
—
Posted
13 days ago
Freshness
Confirmed live yesterday
Closes
Mar 17, 2027

Market check

Salary context

How this pay compares to similar roles

Similar $203k
$127k most similar roles pay here $291k

This listing doesn't post a salary. Most similar roles pay $174,600–$231,000.

Based on 240 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 518 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 505 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Principal AI Software Engineer

The Principal AI Software Engineer joins the Compute System Architecture team within the SPARC organization to drive hardware and software co-design for Azure infrastructure. This role focuses on developing full system software prototypes and evaluation frameworks to optimize memory TCO through tiering, pooling, and overcommit solutions. The engineer will lead the characterization of Large Language Model inference workloads, specifically targeting KV Cache management across GPU HBM, host DRAM, CXL memory, and SSD-backed tiers. Key responsibilities include analyzing data movement across GPU, CPU, storage, and networking subsystems using technologies like CUDA, NCCL, GPUDirect Storage, and GPUDirect RDMA. The role requires proficiency in C, C++, Python, and inference frameworks such as vLLM and TensorRT-LLM to solve complex problems regarding memory hierarchy innovations and distributed inference pipelines for large-scale AI deployments.

What you'll do

  • Prototype full system software to evaluate hardware/software co-designed capabilities for memory TCO reduction.
  • Characterize and optimize Large Language Model (LLM) inference workloads focusing on KV Cache management across various memory tiers.
  • Develop proof-of-concepts and evaluation frameworks to assess memory-tiering architectures for AI inference.
  • Execute workload characterization studies for agentic, multi-turn, and long-context AI workloads to quantify performance metrics.
  • Analyze end-to-end data movement across GPU, CPU, storage, and networking subsystems to identify optimization opportunities.
  • Develop software prototypes and instrumentation to evaluate KV Cache offload, prefetching, and memory-overcommit techniques.
  • Build performance models and simulation frameworks to predict the impact of memory hierarchy innovations on large-scale deployments.
  • Influence hardware architecture and industry alignment through data-driven analysis and technical recommendations.

What we're looking for

  • Bachelor's Degree in Computer Science or related technical field.
  • 6+ years of technical engineering experience coding in languages such as C, C++, C#, Java, JavaScript, or Python.
  • Ability to pass Microsoft Cloud Background Check and other required security screenings.
  • 12+ years of experience in systems software including OS kernel, memory management, I/O stacks, and virtualization (preferred).
  • 10+ years of experience leading hardware/software co-design projects involving CPU or systems architecture (preferred).
  • Deep expertise in Linux kernel internals, memory management, I/O subsystems, NUMA, DMA, and GPU/CPU/storage/network data paths (preferred).
  • Hands-on experience with NVIDIA GPU software stacks including CUDA, NCCL, GDS, and GDR (preferred).
  • Experience with AI inference infrastructure, large GPU clusters, and optimization of KV Cache intensive workloads (preferred).

More like this

Similar roles

Principal AI Software Engineer

Palo Alto Networks

Santa Clara, CA 111 days ago $223,000–$306,500
Generative AI LLMs Python Multi-Agent Systems GPT-4 Llama 3 Anthropic LangSmith Arize HoneyHive OpenTelemetry Java Node.js Agile Scrum API Design
10+ yrs exp

Principal AI Software Engineer

Amd

San Jose, CA 162 days ago $240,000–$360,000
ROCm CUDA OpenCL GPU Architecture CPU Architecture Diffusion MoE Compiler Operating Systems Performance Engineering Hardware/Software Co-design AI/ML Frameworks Open Source

Principal Software Engineer, Core AI Platform

JPMorgan Chase

Seattle, WA 31 days ago
Python AI Infrastructure Machine Learning Cloud Platforms Distributed Systems SRE Observability System Design SDKs Automated Testing Infrastructure Engineering Generative AI
10+ yrs exp

Principal AI Engineer

Humana

New York, NY +2 11 days ago $206,600–$284,300
LLMs RAG Python TypeScript JavaScript React Next.js PostgreSQL Vertex AI Gemini Docker Kubernetes CI/CD Agentic Workflows Model Context Protocol (MCP) OCR Document AI Distributed Systems Cloud-native Architecture
10+ yrs exp Hybrid