Principal Software Engineer, ML & Distributed Systems

Microsoft

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Redmond, WA
Salary
$142,800–$274,800 / yr
Posted
10 days ago
Freshness
Confirmed live 2 days ago
Closes
Feb 28, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $209k
This role $209k
$127k most similar roles pay here $291k

This role pays more than 55% of similar roles. Most pay $174,600–$242,500 — the shaded band above. At the midpoint, this role pays about $209k versus about $209k for comparable roles.

Based on 240 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 598 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 586 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Principal Software Engineer, ML & Distributed Systems

As a Principal Software Engineer, ML & Distributed Systems, you will join the team responsible for powering AI experiences through high-scale distributed services. You will own systems end-to-end by designing, building, and operationalizing scalable machine learning and deep learning models using containers and orchestration platforms like Kubernetes. Your daily work involves refining LLM prompt and fine-tuning strategies, building evaluation pipelines, and managing model serving infrastructure including caching, batching, and GPU capacity management. You will also architect multi-tiered distributed services, develop data pipelines, and manage feature stores to ensure reliable data flow. The role requires proficiency in languages such as C, C++, C#, Java, JavaScript, or Python. You will solve complex problems regarding content moderation and personalization by building high-availability systems that optimize for quality, latency, and cost across the ML and systems boundary.

What does a Software Engineer earn in Washington?

Median $202800 from 326 postings across 28 companies.

See salary data

What you'll do

  • Design, build, and operationalize scalable machine learning and deep learning models using containers and orchestration platforms.
  • Develop and refine LLM prompt engineering and fine-tuning strategies while building evaluation pipelines to optimize quality, latency, and cost.
  • Architect and operate multi-tiered distributed services at hyperscale with high availability and fault tolerance.
  • Build model serving and inference infrastructure including caching, batching, GPU capacity management, and A/B experimentation.
  • Design scalable APIs, data pipelines, and feature stores to ensure reliable data flow between ML systems and products.
  • Drive live-site excellence through instrumentation, monitoring, capacity planning, and incident response for ML-backed services.
  • Mentor engineers across both machine learning and systems disciplines while participating in architectural discussions.

What we're looking for

  • Bachelor's degree in Computer Science or a related technical field is required.
  • Candidates must have 6+ years of technical engineering experience with coding in languages like C, C++, C#, Java, JavaScript, or Python.
  • A Master's degree and 8+ years of experience, or a Bachelor's degree and 12+ years of experience, are preferred qualifications.
  • Experience is required in designing, developing, and operating multi-tiered distributed services at scale.
  • Candidates must have hands-on experience building, deploying, and operating machine learning systems in production.
  • Proficiency in LLM application patterns including prompt engineering, RAG, fine-tuning, and evaluation frameworks is preferred.
  • Experience with ML infrastructure such as distributed training, inference optimization, GPU management, and Kubernetes is preferred.
  • Experience with large-scale data systems including streaming, caching (e.g., Redis), feature stores, and experimentation platforms is preferred.

More like this

Similar roles

Principal AI/ML Engineer

General Motors (GM)

Sunnyvale, CA +5 74 days ago $296,300$423,900
Machine Learning Robotics Python C++ Trajectory Planning Motion Planning Reinforcement Learning Imitation Learning Generative Models Distributed ML Pipelines Simulation Embodied AI Software Engineering
Hybrid

Principal Software Engineer, Distributed Systems Engineer

Nvidia

Remote (Durham, NC) 78 days ago $272,000$431,250
Kubernetes GPU Go Python Slurm Bright Cluster Manager Distributed Systems Cluster Management Monitoring Data Structures Algorithms Systems Programming Network Telemetry Incident Management
10+ yrs exp Remote

Principal Machine Learning Engineer

General Motors (GM)

Remote (Sunnyvale, CA) +5 176 days ago $296,300$453,200
Python PyTorch TensorFlow C++ Distributed Training DDP FSDP GPU Computing AWS GCP Azure Data Processing Pipelines Model Optimization Profiling Debugging
8+ yrs exp Remote

Principal Machine Learning Engineer

Oracle

Seattle, WA +3 53 days ago
Machine Learning PyTorch TensorFlow Keras ETL Version Control Software Development Debugging Infrastructure Data Quality Scalability
6+ yrs exp