Principal Cloud & AI Workload Performance Analysis Engineer

Amd

Confirmed live today High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Austin, TX
Salary
$200,000–$300,000 / yr
Posted
2 days ago
Freshness
Confirmed live today
Closes
Sep 11, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $206k
This role $250k
$132k most similar roles pay here $318k

This role pays more than 82% of similar roles. Most pay $169,550–$241,750 — the shaded band above. At the midpoint, this role pays about $250k versus about $206k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 379 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 379 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Principal Cloud & AI Workload Performance Analysis Engineer

As a Principal Cloud & AI Workload Performance Analysis Engineer, you will join the team to own workloads exposed to hyperscaler custom CPUs, Arm ecosystem momentum, and AI-system architecture. You will measure cloud-native performance, price-performance, and software maturity while analyzing host-CPU effects on accelerator utilization. Your daily responsibilities include defining workload taxonomies, designing reproducible methods for microservices and data processing, and comparing AMD systems against Intel and Arm platforms. You will analyze memory bandwidth, NUMA behavior, and interconnect bottlenecks to translate scaling trends into long-term forecasts. The role requires expertise in Linux systems, container orchestration, hardware performance counters, and system telemetry. You will use these skills to solve complex problems regarding CPU-to-accelerator data movement and cross-ISA deployability while producing technical reports for partners and internal stakeholders within the cloud and AI infrastructure domain.

What you'll do

  • Define and maintain the cloud-native and AI workload taxonomy used for competitive analysis and forecasting.
  • Design reproducible benchmarking methods for web services, microservices, containers, and scale-out data processing.
  • Compare AMD, Intel, Arm, and cloud-custom environments using aligned software versions and tuning policies.
  • Analyze host-CPU impacts on accelerator utilization, memory movement, NUMA behavior, and end-to-end AI system throughput.
  • Measure and model performance-per-watt, density, and price-performance metrics across various platforms.
  • Track compiler, kernel, library, and orchestration factors affecting Arm and cross-ISA deployability.
  • Translate scaling behavior and software trends into 24-36 month forecast inputs and risk assessments.
  • Produce decision-ready technical and executive reports for internal stakeholders and external partners.

What we're looking for

  • Bachelor's or Master's degree in electrical or computer engineering.
  • Deep hands-on experience benchmarking and profiling cloud, distributed, or accelerated-system workloads.
  • Strong Linux systems skills with proficiency in automation or scripting for deployment and analysis.
  • Experience with containers, orchestration, and cloud-native workloads at scale.
  • Ability to design fair experiments across dissimilar CPU architectures and platform configurations.
  • Experience with application profiling, hardware performance counters, and system telemetry.
  • Understanding of AI host-side bottlenecks, NUMA, memory bandwidth, and interconnect effects.
  • Strong technical writing and presentation skills to collaborate with external partners and produce executive reports.
  • Hands-on experience with Arm Neoverse or cloud-custom Arm platforms (preferred).
  • Experience benchmarking major cloud-service-provider instances and services (preferred).
  • Experience measuring AI system throughput or accelerator utilization as a function of host-CPU behavior (preferred).
  • Familiarity with LPDDR/HBM systems, coherent CPU-accelerator links, or high-speed networking (preferred).
  • Participation in MLCommons, OCP, or other relevant benchmark and standards communities (preferred).

More like this

Similar roles

Lead CPU Performance Analysis Engineer

Qualcomm

Santa Clara, CA +2 25 days ago $273,000$409,400
C C++ Python Bash Perl x86 Linux Android Windows GCC LLVM Perf Vtune DynamoRio PIN Memcached NGINX Redis DCPerf
8+ yrs exp

AI Systems Performance Engineer

Broadcom

San Jose, CA 147 days ago $141,300$226,000
Ethernet MLPerf NCCL Python C++ Linux PyTorch RDMA RoCEv2 Docker Kubernetes CI/CD Performance Benchmarking Distributed Systems
10+ yrs exp

Principal Software Engineer, Performance

Microsoft

13 days ago $165,600$296,400
C++ Python AI Frameworks Performance Engineering Benchmarking Azure Compiler Computer Architecture Observability Software Architecture
8+ yrs exp