Capacity & Efficiency Infrastructure

Microsoft

Confirmed live 2 days ago High trust
Closes in 5 days Hybrid

Quick summary

Work type
Hybrid
Location
Salary
$119,800–$234,700 / yr
Posted
175 days ago
Freshness
Confirmed live 2 days ago
Closes
Sep 16, 2026 (soon)

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $179k
This role $177k
$100k most similar roles pay here $249k

This role pays more than 62% of similar roles. Most pay $144,350–$213,962 — the shaded band above. At the midpoint, this role pays about $177k versus about $179k for comparable roles.

Based on 240 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 598 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 586 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Capacity & Efficiency Infrastructure

As a Member of Technical Staff – Capacity & Efficiency Infrastructure within the Microsoft AI team, you will focus on improving the management and efficiency of a large-scale compute fleet. You will develop software and mathematical models to measure capacity usage while building tools and techniques to optimize performance across thousands of GPUs in a supercomputing environment. Your daily work involves designing and optimizing distributed training infrastructure, developing telemetry systems for visibility into model performance, and profiling bottlenecks across memory, networking, and storage subsystems. You will utilize Python, C++, CUDA, Triton, NCCL, PyTorch, and JAX to solve complex problems related to large-scale machine learning and generative AI workloads. Key responsibilities include implementing distributed training parallelism, optimizing collective communication libraries for NVLink and InfiniBand topologies, and collaborating with researchers to scale advanced research recipes.

What you'll do

  • Design, implement, and optimize distributed training infrastructure in Python and C++ for large-scale GPU clusters.
  • Build telemetry systems to monitor infrastructure performance, ML model utilization, and cost metrics.
  • Profile, benchmark, and debug performance bottlenecks across compute, memory, networking, and storage subsystems.
  • Drive architectural improvements across machine learning services to achieve measurable efficiency gains.
  • Develop automated tools that provide insights and recommendations for improving fleet-wide efficiency.
  • Optimize collective communication libraries like NCCL for emerging NVLink and InfiniBand topologies.
  • Partner with researchers and engineers to balance infrastructure growth with operational efficiency.
  • Collaborate with hardware teams to optimize software for next-generation accelerators.

What we're looking for

  • Bachelor's degree in Computer Science or a related technical discipline is required.
  • At least 6 years of technical engineering experience coding in languages such as C, C++, C#, Java, JavaScript, or Python.
  • Advanced proficiency in C++ and/or Python for high-performance computing environments.
  • Deep understanding of GPU architectures and DL/LLM architectures.
  • Extensive experience profiling and analyzing performance in large-scale distributed computing systems and GenAI models.
  • Experience with low-level GPU programming (CUDA, Triton, NCCL) and frameworks like PyTorch or JAX.
  • Experience building infrastructure for large-scale machine learning or generative AI workloads.
  • Expertise in networking (InfiniBand, NVLink), storage systems, or distributed training parallelisms.

More like this

Similar roles

Multimodal Infrastructure

Microsoft

20 days ago $142,800$274,800
Python C++ C# Java JavaScript PyTorch Megatron Deepspeed Ray Spark vLLM TensorRT-LLM SGLang Triton RLHF DPO GRPO Quantization Distributed Training data processing pipelines
6+ yrs exp Hybrid

Infrastructure Engineering

State Street

Remote (Boston, MA) 66 days ago $178,131$195,944
Terraform Docker Kubernetes Helm AWS Azure OCI Python Bash Jenkins Ansible Git Splunk Data robot Snowflake Databricks Elasticsearch ServiceNow CI/CD IaC
5+ yrs exp Remote

Data Center Capacity Planner, Infrastructure Services

Apple Inc

Sunnyvale, CA 2 days ago $140,700$268,400
Data Center Operations Capacity Planning Cloud Platforms AI Distributed Systems Infrastructure Telemetry Observability Automation Data Analysis Hybrid Infrastructure Infrastructure Strategy Capacity Governance
3+ yrs exp

Data Center Capacity Planner, Infrastructure Services

Apple Inc

Austin, TX 2 days ago
Data Center Operations Capacity Planning Cloud Platforms AI Infrastructure Planning Distributed Systems Telemetry Observability Automation Infrastructure Strategy Hybrid Infrastructure Governance
3+ yrs exp

Infrastructure Engineers

Shopify

Remote 134 days ago
Kubernetes Docker Terraform Prometheus Kafka MySQL Redis GCP infrastructure-as-code Distributed Systems Monitoring Observability
Remote