Software Development Engineer, GPU Fleet Management & AI Infrastructure

Amd

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CA
Salary
$204,000–$306,000 / yr
Posted
5 days ago
Freshness
Confirmed live yesterday
Closes
Sep 10, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $209k
This role $255k
$149k most similar roles pay here $323k

This role pays more than 86% of similar roles. Most pay $177,237–$241,562 — the shaded band above. At the midpoint, this role pays about $255k versus about $209k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 379 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 379 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Software Development Engineer, GPU Fleet Management & AI Infrastructure

Software Development Engineer — GPU Fleet Management & AI Infrastructure will join a team at the intersection of systems software, cloud infrastructure, and accelerated computing to build Fleet Manager. This role involves developing a secure control plane for large-scale AMD GPU infrastructure, focusing on scheduling, workload orchestration, hardware health management, and inference capabilities. You will design distributed control-plane services, command-line tools, and multi-tenant infrastructure while integrating with Kubernetes, Kueue, JobSet, and ROCm. The position requires expertise in Rust, C++, or Go to build reliable systems involving gRPC, REST, and PostgreSQL. Key responsibilities include managing GPU training workloads, implementing admission control, and ensuring security through authentication and isolation. You will solve complex infrastructure problems related to distributed system failure modes, hardware telemetry, and the deployment of large-scale AI models across diverse execution environments like Slurm or custom container runtimes.

What does a Software Development Engineer earn in California?

Median $217725 from 74 postings across 8 companies.

See salary data

What you'll do

  • Design and develop distributed control-plane services, APIs, schedulers, and command-line tools for GPU infrastructure.
  • Build reliable orchestration systems for GPU training, inference, and interactive development workloads.
  • Develop scalable scheduling capabilities including topology-aware placement, priority, and multi-node workload coordination.
  • Implement durable reconciliation, lifecycle management, and recovery mechanisms across PostgreSQL and external execution systems.
  • Integrate Fleet Manager with Kubernetes and related technologies like Kueue, JobSet, and various storage systems.
  • Create GPU health diagnostics, quarantine capabilities, and remediation tools using ROCm and hardware telemetry.
  • Build secure multi-tenant infrastructure featuring robust authentication, authorization, and workload isolation.
  • Improve the performance of AI inference services through routing, streaming, load shedding, and usage metering.

What we're looking for

  • Experience in systems software development using Rust, C++, Go, or comparable languages (preferred).
  • Production experience with the Rust programming language is highly desirable (preferred).
  • Experience designing and operating distributed systems, control planes, schedulers, or cloud infrastructure (preferred).
  • Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery (preferred).
  • Experience building reliable services using REST, streaming, WebSocket, or gRPC APIs (preferred).
  • Experience with Kubernetes internals, controllers, operators, scheduling, resource management, or custom resources (preferred).
  • Experience with PostgreSQL-backed services, schema evolution, transactions, leader election, and optimistic concurrency (preferred).
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience (preferred).

More like this

Similar roles

Software Engineer, GPU AI ML

Amd

Santa Clara, CA 82 days ago $204,000$306,000
C++ HIP CUDA ROCm PyTorch TensorFlow JAX GPU Architecture Kernel Optimization Distributed Systems LLMs SFT RLHF GRPO Quantization Verilog SystemVerilog RTL Design
Hybrid

Senior Datacenter GPU Firmware Engineer

Amd

Santa Clara, CA +1 7 days ago $161,200$241,800
Firmware GPU Drivers OS Kernel C C++ Python Bash Linux CUDA ROCm OpenCL Hardware Bring-up Server Architecture BIOS Performance Optimization
Hybrid

Senior Software Engineer, GPU Local AI Platforms

Nvidia

Santa Clara, CA +4 54 days ago $224,000$356,500
Python C++ CUDA Triton LLM Inference GPU Computing Docker CI/CD NCCL RCCL Quantization Tensor Parallelism Kernel Development NVIDIA Container Toolkit
10+ yrs exp