Senior Software Engineer, GPU Local AI Platforms

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CAWestford, MAAustin, TXDurham, NCSeattle, WA
Salary
$224,000–$356,500 / yr
Posted
49 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $193k
This role $290k
$123k most similar roles pay here $382k

This role pays more than 99% of similar roles. Most pay $151,000–$235,750 — the shaded band above. At the midpoint, this role pays about $290k versus about $193k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, GPU Local AI Platforms

As a Senior Software Engineer - GPU Local AI Platforms, you will join the Local AI team to build the software stack enabling large language models and generative AI applications to run efficiently on edge hardware. You will track innovations in open-source LLM inference frameworks, analyze how new architectures like MoE routing or speculative decoding map onto GPU architecture, and characterize multi-node inference behavior using NCCL/RCCL. Your daily work involves producing performance analysis reports, owning model validation workflows, and maintaining developer-facing inference recipes. The role requires expertise in Python, C++, CUDA, Triton, and container engineering with Docker. You will solve technical challenges regarding memory bandwidth, throughput, and latency while serving as a technical point of contact for hardware-specific issues to ensure community innovations are reliable at scale for developers and partners.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Evaluate innovations in open-source LLM inference frameworks to identify performance-critical features for edge hardware.
  • Analyze how new model architectures and inference algorithms map onto NVIDIA GPU architecture for optimization.
  • Characterize multi-node inference behavior including collective communication primitives and parallelism efficiency on edge clusters.
  • Produce performance analysis reports mapping theoretical hardware limits to observed inference throughput, latency, and utilization.
  • Manage the model validation workflow for new releases, including compatibility assessment and recipe development.
  • Develop and maintain accurate developer-facing inference recipes while automating staleness detection from CI results.
  • Serve as the technical point of contact for partners regarding hardware-specific inference issues.

What we're looking for

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • 12+ years of software engineering experience with depth in GPU computing, ML systems, or high-performance inference.
  • Strong Python or C++ programming skills and expertise in software design and engineering.
  • Hands-on experience with GPU kernel development or optimization using CUDA, C++, Triton, or equivalent.
  • Working knowledge of LLM inference internals including attention mechanisms, KV-cache management, continuous batching, quantization, and tensor parallelism.
  • Expertise in container engineering including multi-architecture Docker/OCI builds and NVIDIA Container Toolkit.
  • Strong analytical skills to form performance hypotheses, design experiments, and communicate findings clearly.

More like this

Similar roles

Senior Deep Learning Frameworks CUDA Software Engineer

Nvidia

Remote (Santa Clara, CA) +1 11 days ago $184,000$287,500
CUDA PyTorch JAX C++ Python TRT-LLM vLLM SGLang TensorRT Triton NCCL MPI UCX XLA HPC Kernel Authoring NVIDIA Nsight Systems Compiler Technologies
8+ yrs exp Remote

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Senior Software Engineer, AI Inference Performance

Nvidia

Santa Clara, CA 16 days ago $184,000$287,500
CUDA Python C++ Rust TensorRT-LLM vLLM SGLang Triton PyTorch NVIDIA Nsight Systems CUTLASS NCCL Quantization Distributed Systems GPU Architecture Model Parallelism Speculative Decoding
6+ yrs exp

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Senior Software Engineer, AI Platform

Anduril Industries

Seattle, WA 60 days ago $191,000$253,000
CI/CD Python C++ Rust Terraform Ansible Docker Kubernetes GitLab CI Jenkins GitHub Actions Crossplane Puppet Chef Bash Prometheus Grafana ELK Stack Splunk GitOps