Principal Software Engineer, AI Performance & Reliability

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CABellevue, WA
Salary
$240,000–$360,000 / yr
Posted
2 days ago
Freshness
Confirmed live yesterday
Closes
Oct 1, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $208k
This role $300k
$117k most similar roles pay here $386k

This role pays more than 90% of similar roles. Most pay $177,000–$238,334 — the shaded band above. At the midpoint, this role pays about $300k versus about $208k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 498 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 498 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Principal Software Engineer, AI Performance & Reliability

Principal Software Engineer — AI Performance & Reliability will join the AI Infrastructure team to improve the performance, efficiency, and reliability of machine learning workloads across training and inference. This role focuses on optimizing large language models, diffusion models, and recommendation systems by identifying bottlenecks across frameworks, compilers, runtimes, operating systems, and hardware. The engineer will develop performance tooling, benchmarks, and observability systems while collaborating with customers to resolve complex production issues and translate feedback into infrastructure improvements. Required skills include proficiency in Python or C++, experience with PyTorch, TensorFlow, or JAX, and a strong foundation in computer architecture and distributed communication. Candidates should have expertise in GPU environments and tools like ROCm, HIP, CUDA, Triton, XLA, MLIR, or NCCL to solve technical challenges across the AI software and hardware stack.

What does a Software Engineer earn in California?

Median $214000 from 869 postings across 70 companies.

See salary data

What you'll do

  • Profile and optimize performance, efficiency, and reliability for AI model training and inference workloads.
  • Identify and resolve technical bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
  • Improve model throughput, latency, memory efficiency, and scalability for large language models and diffusion models.
  • Develop performance tooling, benchmarks, automation, and observability systems to monitor AI infrastructure.
  • Investigate and resolve complex production issues affecting high-scale machine learning workloads.
  • Partner directly with customers to understand technical requirements, reproduce issues, and provide tailored solutions.
  • Translate customer feedback into actionable product improvements and infrastructure capabilities.
  • Document performance findings, technical recommendations, and best practices for internal teams.

What we're looking for

  • A PhD or a master's degree in artificial intelligence, machine learning, computer science, or a related field.
  • Strong software engineering skills and experience building production-quality systems.
  • Experience working with AI infrastructure for model training, inference, or both.
  • Demonstrated experience profiling and optimizing machine learning models or AI workloads.
  • Strong foundations in computer architecture including processors, memory hierarchies, parallelism, and performance tradeoffs.
  • Proficiency in systems-oriented programming languages such as Python or C++.
  • Experience with machine learning frameworks like PyTorch, TensorFlow, or JAX.
  • Experience with GPU, accelerator, distributed computing, or specific tools like ROCm, HIP, CUDA, Triton, XLA, MLIR, and NCCL (preferred).

More like this

Similar roles

AI Engineer, Recursive Self-Improvement for Compute

Amd

Santa Clara, CA 67 days ago $204,000–$306,000
Python C++ CUDA HIP ROCm Triton PyTorch JAX TensorFlow Reinforcement Learning GPU Kernels Compiler Optimization Performance Engineering Profiling Distributed Training Agentic Workflows Hardware-aware Optimization
Hybrid

Principal Software Engineer, Core AI Platform

JPMorgan Chase

Seattle, WA 33 days ago
Python AI Infrastructure Machine Learning Cloud Platforms Distributed Systems SRE Observability System Design SDKs Automated Testing Infrastructure Engineering Generative AI
10+ yrs exp

Principal AI Software Engineer

Palo Alto Networks

Santa Clara, CA 113 days ago $223,000–$306,500
Generative AI LLMs Python Multi-Agent Systems GPT-4 Llama 3 Anthropic LangSmith Arize HoneyHive OpenTelemetry Java Node.js Agile Scrum API Design
10+ yrs exp

Principal AI Software Engineer

Amd

San Jose, CA 164 days ago $240,000–$360,000
ROCm CUDA OpenCL GPU Architecture CPU Architecture Diffusion MoE Compiler Operating Systems Performance Engineering Hardware/Software Co-design AI/ML Frameworks Open Source

Principal AI Software Engineer

Microsoft

15 days ago $142,800–$274,800
C C++ Python CUDA Linux Kernel Memory Management NVIDIA GPU NCCL GPUDirect Storage GPUDirect RDMA vLLM SGLang TensorRT-LLM CXL Distributed Systems Virtualization NUMA DMA I/O Stacks
6+ yrs exp Hybrid

Principal AI Software Engineer

Palo Alto Networks

Santa Clara, CA 2 days ago $223,000–$306,500
Generative AI LLM Python Multi-Agent Systems GPT-4 Llama 3 Anthropic LangSmith Arize HoneyHive OpenTelemetry Java Node.js Agile Scrum API Design
10+ yrs exp