Cloud & Customer Solutions Engineer, DC GPU

Amd

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Bellevue, WA
Salary
$204,000–$306,000 / yr
Posted
15 days ago
Freshness
Confirmed live 2 days ago
Closes
Aug 27, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $194k
This role $255k
$124k most similar roles pay here $325k

This role pays more than 89% of similar roles. Most pay $152,000–$235,750 — the shaded band above. At the midpoint, this role pays about $255k versus about $194k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Cloud & Customer Solutions Engineer, DC GPU

Cloud & Customer Solutions Engineer - DC GPU joins the Data Center GPU organization on the Applied AI team to support strategic customers like frontier labs and cloud providers. This role involves managing the end-to-end lifecycle of AMD Instinct GPU clusters, including bring-up, certification, workload deployment, and production incident response. You will build and tune large-scale training and inference stacks while developing observability and benchmarking tools to ensure production readiness. The position requires technical expertise in ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm, Python, and systems languages. You will also work with high-performance networking like RoCE or InfiniBand across cloud and bare-metal environments. The role focuses on solving complex infrastructure challenges for AI workloads, ensuring stable performance in distributed training and inference while translating field findings into upstream improvements for the ROCm ecosystem and reference architectures.

What you'll do

  • Manage end-to-end customer deployments including cluster bring-up, certification, workload onboarding, and performance validation.
  • Deploy and tune large-scale training and inference stacks across cloud and bare-metal environments.
  • Lead root-cause analysis and resolution for production incidents on customer GPU clusters.
  • Build observability, benchmarking, and validation tools to certify clusters as production-ready.
  • Transfer operational capabilities to customers through documentation, runbooks, and hands-on enablement.
  • Convert field findings into upstream contributions for ROCm, serving frameworks, and reference architectures.
  • Deploy agentic AI solutions and manage their production behavior within customer environments.

What we're looking for

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • 5+ years of production software or infrastructure engineering experience.
  • Hands-on experience with GPU compute at scale, including cluster deployment and distributed training.
  • Strong knowledge of AI infrastructure tools like Kubernetes, Slurm, ROCm, vLLM, SGLang, and RCCL.
  • Proficiency in Python and at least one systems language to navigate large codebases.
  • Experience with high-performance networking (RoCE/InfiniBand) and observability tooling (Prometheus, Grafana).
  • Direct customer-facing experience involving technical escalations or embedded deployments.
  • Must not require visa sponsorship to work in the United States.

More like this

Similar roles

Senior Datacenter GPU Firmware Engineer

Amd

Santa Clara, CA +1 3 days ago $161,200$241,800
Firmware GPU Drivers OS Kernel C C++ Python Bash Linux CUDA ROCm OpenCL Hardware Bring-up Server Architecture BIOS Performance Optimization
Hybrid

Senior Cloud Solutions Engineer

General Dynamics

Silver Spring, MD +1 14 days ago $129,813$166,750
Microsoft Azure Artificial Intelligence Large Language Models DevOps Infrastructure as Code Terraform Python PowerShell TypeScript JavaScript CI/CD Retrieval-Augmented Generation Vector Search Entra ID Git Cloud Networking
8+ yrs exp

Lead Systems Debug Engineer, Data Center GPU

Amd

Austin, TX 64 days ago $163,200$244,800
GPU SoC PCIe HBM C C++ Python Shell Perl Git Agile Root-Cause Analysis Oscilloscopes Board-level Diagnostics Power Delivery Networking Cloud Infrastructure
8+ yrs exp Hybrid

Senior Cloud Software Engineer

Nvidia

Remote 6 days ago $152,000$241,500
Kubernetes AWS GCP Azure Go Python Rust C++ Java Distributed Systems Cloud-Native Data Management Storage Systems Performance Engineering Observability
5+ yrs exp Remote