Post-Training Platform Infrastructure Engineer

Amd

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CA
Salary
$204,000–$306,000 / yr
Posted
87 days ago
Freshness
Confirmed live 2 days ago
Closes
Jun 16, 2027

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $177k
This role $255k
$83k most similar roles pay here $330k

This role pays more than 94% of similar roles. Most pay $151,000–$203,723 — the shaded band above. At the midpoint, this role pays about $255k versus about $177k for comparable roles.

Based on 240 similar postings.

Employer

About Amd

AMD (Advanced Micro Devices) is a semiconductor company that develops high-performance processors, graphics cards, and adaptive computing solutions for gaming, data centers, and embedded markets. Industry: Semiconductors

Amd currently has 367 open roles on FindRole.

Listed pay typically runs $166,400–$249,600 across 367 roles with salary data.

Most-posted roles

View all roles at Amd

At a glance

TL;DR · Post-Training Platform Infrastructure Engineer

As a Post-Training Platform Infrastructure Engineer, you will join the team to focus on post-training and inference infrastructure, specifically addressing P/D disaggregation, KV cache lifecycle management, and efficient offloading mechanisms. You will research and analyze modern LLM inference frameworks to identify performance bottlenecks in latency, throughput, and memory pressure while developing infrastructure-level features for multi-GPU and multi-node deployments. The role also involves applying a research-driven approach to RL systems to optimize policy rollout and memory usage during training. To succeed, you must possess expertise in distributed systems, GPU-accelerated workloads, and memory-constrained environments. You will utilize Python and C++ to modify complex open-source codebases while leveraging profiling tools for CPU and GPU performance analysis. Your work solves critical challenges in large-scale model inference and reinforcement learning pipelines.

What you'll do

  • Research and analyze architectural tradeoffs of prefill/decode disaggregation in LLM inference frameworks.
  • Manage KV cache lifecycles, including memory layout, eviction strategies, and reuse across various backends.
  • Develop infrastructure features to optimize inference latency, throughput, and memory efficiency.
  • Implement offloading mechanisms for KV caches across GPU, CPU, and storage systems.
  • Enhance scalability of inference and reinforcement learning (RL) systems across multi-GPU and multi-node deployments.
  • Debug performance and correctness issues within distributed RL pipelines and post-training loops.
  • Translate model-level requirements into production-grade infrastructure capabilities through research-driven engineering.
  • Validate performance gains using benchmarks and document architectural insights to guide future system design.

What we're looking for

  • Bachelor's or master's degree in computer science, computer engineering, electrical engineering, or equivalent.
  • Strong background in systems engineering, distributed systems, or ML infrastructure.
  • Hands-on experience with GPU-accelerated workloads and memory-constrained systems.
  • Proficiency in Python and C++ or similar systems languages.
  • Experience debugging performance issues using CPU, GPU, and memory profiling tools.
  • Ability to read, understand, and modify complex open-source codebases.
  • Knowledge of LLM inference workflows, attention mechanisms, and KV cache management.
  • Experience with distributed RL pipelines or post-training infrastructure.

More like this

Similar roles

Senior Software Engineer, RL Post-Training Frameworks

Nvidia

Remote (Santa Clara, CA) 21 days ago $184,000$287,500
Reinforcement Learning Python C/C++ PyTorch Kubernetes Ray vLLM SGLang TensorRT-LLM DeepSpeed Megatron-LM NCCL InfiniBand FSDP Distributed Systems High-Performance Computing
5+ yrs exp Remote

Senior Staff LLM Serving Engineer, Cloud AI Engineering

Qualcomm

San Diego, CA +1 155 days ago $158,400$237,600
LLM PyTorch Python Triton-Inference Server vLLM SGLang CUDA Triton torch.compile torchDynamo Distributed Systems Kernel Design Deep Learning KV-Cache Management Model Optimization
4+ yrs exp

Infrastructure Platform Engineer

Cisco

Remote 21 days ago $116,600$147,900
VMware Linux Windows Server Ansible Terraform PowerShell DSC TCP/IP DNS DHCP Active Directory ServiceNow Infrastructure Monitoring Configuration Management Networking
5+ yrs exp Remote

Platform Engineer

Cisco

Remote 29 days ago $139,300$203,600
CI/CD Linux Ubuntu RHEL CentOS Infrastructure as Code Puppet Ansible Kubernetes Python Go Proxmox Jenkins Slurm Ceph Elasticsearch Nix NixOS Bare-metal
7+ yrs exp Remote

Platform Engineer

Booz Allen Hamilton

Fort Belvoir, VA 43 days ago $62,000$141,000
Kubernetes AWS Azure GCP Terraform Ansible CloudFormation Puppet CI/CD GitLab GitHub Python Bash PowerShell Groovy Ruby JSON REST XML YAML Infrastructure-as-Code

Platform Engineer

Apex

Austin, TX 43 days ago
GCP Terraform infrastructure-as-code CI/CD GitHub Actions Kubernetes GKE Docker Cloud Run Cloud SQL Pub/Sub BigQuery Python Bash Go Apache Airflow Prometheus Grafana Datadog FinOps
3+ yrs exp Hybrid