Quick summary
- Work type
- Hybrid
- Location
- —
- Salary
- $142,800–$274,800 / yr
- Posted
- 32 days ago
- Freshness
- Confirmed live 2 days ago
- Closes
- Feb 20, 2027
Employer
About Microsoft
Microsoft Corporation is a global technology leader producing software, hardware, and cloud services including Windows, Office 365, Azure cloud platform, Xbox gaming, and Surface devices. Industry: Software & Cloud Computing
Microsoft currently has 644 open roles on FindRole.
Listed pay typically runs $119,800–$234,700 across 624 roles with salary data.
Most-posted roles
- Software Engineer 216
- Product Manager 29
- Technical Program Manager 19
- Applied Scientist 15
- Cloud Solution Architect 13
At a glance
TL;DR · Principal AI Accelerator Tools Development Engineer
The Principal AI Accelerator Tools Development Engineer joins the Platform Systems Engineering team to lead the development of next-generation stress, validation, and performance tooling for MAIA AI accelerator platforms. This role involves building software frameworks, stress workloads, and validation tools that exercise every layer of the AI stack, including hardware execution engines, memory subsystems, compiler-generated kernels, and distributed communication fabrics. The engineer will develop scalable infrastructure for workload deployment, telemetry collection, and result analysis while optimizing kernels for custom accelerators. Key responsibilities include characterizing system performance across compute, networking, and storage subsystems to ensure platform readiness. The role requires expertise in PyTorch, Triton, Python, C++, and custom MAIA SDKs. The work focuses on solving complex hardware-software integration challenges, ensuring the reliability of large-scale AI training and inference workloads within a high-performance computing environment.
Skills
What you'll do
- Design and develop scalable stress, performance, and validation frameworks for MAIA AI accelerator platforms.
- Build workload generation infrastructure to exercise compute, memory, interconnect, networking, and storage resources.
- Develop and optimize kernels for custom AI accelerators using PyTorch, Triton, Python, and C++.
- Create synthetic and production-inspired workloads that model large-scale training and inference behaviors.
- Integrate tooling with MAIA compiler pipelines, SDKs, and runtime environments to automate validation and testing.
- Characterize system performance across hardware subsystems and develop benchmarking methodologies and dashboards.
- Build automated infrastructure for workload deployment, telemetry collection, and root-cause analysis of performance bottlenecks.
- Develop tools to identify correctness, thermal, power, and stability issues during platform bring-up and qualification.
What we're looking for
- Master's degree in Engineering or related field and 7+ years of experience, or Bachelor's degree and 8+ years of experience.
- Equivalent experience to the specified educational and experience requirements is acceptable.
- 8+ years of experience developing and optimizing AI training and inference workloads for GPUs, accelerators, or HPC platforms using C++, PyTorch, and Triton.
- 8+ years of experience analyzing and optimizing workloads on AI accelerator, GPU, or HPC platforms including performance profiling and bottleneck analysis.
- 8+ years of experience developing kernels and building automated stress, validation, benchmarking, and reliability frameworks across hardware and software environments.
- Ability to pass the Microsoft Cloud Background Check.
- Experience with AI compiler technologies and kernel generation frameworks like LLVM or MLIR (preferred).
- Experience with large-scale AI models, custom accelerator SDKs, and silicon bring-up activities (preferred).
Related searches