Senior AI/ML Platform Engineer

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CA
Salary
$204,000–$306,000 / yr
Posted
51 days ago
Freshness
Confirmed live yesterday
Closes
Jul 22, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $215k
This role $255k
$162k most similar roles pay here $321k

This role pays more than 85% of similar roles. Most pay $177,500–$251,875 — the shaded band above. At the midpoint, this role pays about $255k versus about $215k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Senior AI/ML Platform Engineer

As a Sr. AI/ML Platform Engineer, you will join the team to build the platform layer that makes AI-for-engineering workflows scalable, reliable, and reproducible. You will develop infrastructure for distributed training, inference, experiment tracking, and GPU cluster utilization while transforming research workflows into production services, APIs, and automated systems. Your daily work involves building tools for benchmark automation, artifact management, and performance monitoring to support kernel optimization, verification, and simulation. The role requires proficiency in Python and systems languages like C++, Go, or Rust, alongside experience with Kubernetes, Ray, Slurm, Triton, vLLM, SGLang, and ROCm/HIP. You will solve complex problems regarding distributed system reliability and hardware-software tooling to support large-scale agent execution and automated program optimization within the technical domain of semiconductor engineering and high-performance computing infrastructure.

What you'll do

  • Build and operate shared infrastructure for agentic workflows including job scheduling, orchestration, and experiment tracking.
  • Develop reliable systems for distributed training, inference, and large-scale agent rollout across GPU clusters.
  • Create platform services for benchmark execution, correctness checking, profiling, and regression tracking.
  • Maintain artifact management systems for kernels, RTL edits, logs, and formal verification data.
  • Improve GPU cluster utilization through better scheduling, reliability, quota management, and workload isolation.
  • Build dashboards and observability systems to monitor experiment status, resource usage, and failure modes.
  • Productionize research workflows for RL systems, inference systems, and evaluation pipelines.
  • Integrate the platform with compilers, ROCm/HIP tooling, profilers, simulators, and EDA tools.

What we're looking for

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or a related field.
  • Master's degree preferred; PhD is a plus, especially with experience in ML systems, distributed systems, or GPU computing.
  • Proficiency in Python and at least one systems language such as C++, Go, or Rust.
  • Experience building ML platforms, AI infrastructure, distributed systems, workflow orchestration, or GPU cluster infrastructure.
  • Experience with Kubernetes, Ray, Slurm, containerization, CI/CD, or large-scale compute orchestration.
  • Experience with ROCm/HIP, CUDA, PyTorch, JAX, vLLM, SGLang, Triton, or similar ML frameworks and tools.
  • Expertise in building experiment tracking systems, artifact stores, benchmark automation, or developer productivity tools.
  • Familiarity with compiler, profiler, simulator, formal verification, EDA, firmware, or hardware performance workflows is a plus.

More like this

Similar roles

Senior Machine Learning Engineer, AI Platform

Adobe

San Jose, CA 8 days ago $211,800$306,625
Python Go C++ Rust Java Kubernetes Distributed Systems GPU PyTorch FSDP DeepSpeed vLLM TensorRT-LLM Triton Ray Serve Cloud Infrastructure
7+ yrs exp

ML Systems Research Engineer, RL / Inference / Agent Systems

Amd

Santa Clara, CA 44 days ago $204,000$306,000
Python PyTorch JAX TensorFlow Reinforcement Learning RLHF LLM Agents Kubernetes Ray Slurm CUDA ROCm HIP Distributed Systems Data Pipelines Model Serving Compiler Optimization Profiling Hardware Engineering
Hybrid

Leader, Enterprise AI Platforms

Qualcomm

San Diego, CA 179 days ago $198,500$297,700
Kubernetes Python C++ Java AWS GCP Azure MLOps LLMOps Terraform Bicep Helm Docker CI/CD Slurm Triton Inference Server vLLM KServe PyTorch CUDA Ray LangChain CrewAI AutoGen Milvus Pinecone FAISS Airflow Argo MLflow
10+ yrs exp

AI Engineer, Recursive Self-Improvement for Compute

Amd

Santa Clara, CA 45 days ago $204,000$306,000
Python C++ CUDA HIP ROCm Triton PyTorch JAX TensorFlow Reinforcement Learning GPU Kernels Compiler Optimization Performance Engineering Profiling Distributed Training Agentic Workflows Hardware-aware Optimization
Hybrid

Lead AI Platform Engineer

Abbott

Madison, WI 41 days ago $99,300$198,700
Machine Learning Natural Language Processing Generative AI Large Language Models Python RAG Batch and Stream Processing Model Distillation Agile SAFe Cloud Platforms

Senior Engineer, AI Platforms

Qualcomm

San Diego, CA 72 days ago $111,300$166,900
LLM Python C#.NET JavaScript TypeScript React.js Angular Node.js HTML CSS C C++ API Design Vector Search Elasticsearch Docker Kubernetes Multi-agent Systems
2+ yrs exp