Senior Cloud Hardware Storage Engineer

Microsoft

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Salary
$119,800–$234,700 / yr
Posted
3 days ago
Freshness
Confirmed live yesterday
Closes
Mar 14, 2027

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $199k
This role $177k
$105k most similar roles pay here $256k

This role pays less than 62% of similar roles. Most pay $162,000–$235,750 — the shaded band above. At the midpoint, this role pays about $177k versus about $199k for comparable roles.

Based on 240 similar postings.

Employer

About Microsoft

Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing

Microsoft currently has 657 open roles on FindRole.

Listed pay typically runs $119,800–$234,700 across 642 roles with salary data.

Most-posted roles

View all roles at Microsoft

At a glance

TL;DR · Senior Cloud Hardware Storage Engineer

As a Senior Cloud Hardware Storage Engineer within the Azure Memory and Storage Center of Excellence, you will join the Silicon Cloud Hardware Infrastructure Engineering team to scale fault self-healing and failure prediction systems. You will build the underlying platform for these systems, including telemetry pipelines, prediction services, decision logic, and automated repair workflows. Your daily work involves shipping production AI agents featuring prompt design, retrieval over diagnostics data, and robust CI/CD paths. You will also develop SDKs, APIs, and dashboards to improve developer experience while monitoring metrics like prediction precision and false-repair rates. The role requires expertise in Python, C#, C++, or Rust, along with experience in LLM application patterns, RAG, and high-volume data pipelines. This position addresses the critical challenge of maintaining fleet availability by automating hardware failure detection and remediation across millions of nodes.

What you'll do

  • Build telemetry pipelines, prediction services, and automated repair workflows for Azure's fault self-healing systems.
  • Develop and deploy AI agents with custom prompt designs, retrieval logic, and safety guardrails.
  • Implement automated remediation workflows using staged rollouts and blast-radius limits to ensure safe infrastructure actions.
  • Create SDKs, APIs, and dashboards to improve the developer experience for internal engineering teams.
  • Monitor and report on key metrics including prediction precision, false-repair rates, and action success rates.
  • Design and maintain high-volume data pipelines to support large-scale automated hardware management.
  • Establish regression gates and evaluation harnesses to prevent faulty models from reaching production.

What we're looking for

  • A Bachelor's degree in Engineering or a related field and 5+ years of technical experience, or a Master's degree and 3+ years.
  • 8+ years of experience building and operating production software, distributed services, data platforms, or large-scale automation.
  • Proficiency in Python and C# (or C++/Rust) with experience in high-volume cloud-scale data pipelines.
  • Experience with AI/ML in production including model serving, evaluation, versioning, and monitoring.
  • Knowledge of LLM application patterns such as agent/tool-calling, RAG, and structured output.
  • A track record of automation that takes real actions on infrastructure with necessary safety engineering.
  • Ability to pass the Microsoft Cloud Background Check.
  • Experience in firmware/embedded systems, storage device protocols (NVMe/PCIe), or live-site operations at scale (preferred).

More like this

Similar roles

Senior Software Engineer, Cloud Native Storage

Broadcom

Promontory C, CA 122 days ago $141,300$226,000
Kubernetes Go C++ vSAN CSI Distributed Systems Linux Container Runtime SDLC Data Structures Algorithms Persistent Volumes NFS VMFS etcd Infrastructure Orchestration
5+ yrs exp

Senior Storage Platform Engineer

Nvidia

Santa Clara, CA 9 days ago $168,000$270,250
NetApp Pure Storage Cloudian DDN Infrastructure as Code Ansible Terraform CI/CD Python Go GitOps NFS SMB iSCSI NVMe-oF S3 Lustre Prometheus Grafana Splunk Datadog Docker Kubernetes
8+ yrs exp Hybrid

Senior HPC Storage Engineer

Nvidia

Santa Clara, CA +1 2 days ago $184,000$287,500
Distributed Storage HPC Python Bash Docker Enroot Ceph Weka.io Vast Lustre GPFS CUDA NCCL MLPerf NVMe SDN PyTorch TensorFlow Linux RHEL
8+ yrs exp

Software Engineer, Storage

DoorDash, Inc

San Francisco, CA +3 87 days ago $193,800$285,000
Cassandra Redis Kafka Go Java Memcached DynamoDB CockroachDB Distributed Systems NoSQL Temporal Cadence Argo Multi-threading Data Abstraction
2+ yrs exp Hybrid