Senior AI Performance and Efficiency Engineer

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CANew York, NYSeattle, WA
Salary
$152,000–$241,500 / yr
Posted
176 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $204k
This role $197k
$140k most similar roles pay here $265k

This role pays less than 56% of similar roles. Most pay $162,000–$246,150 — the shaded band above. At the midpoint, this role pays about $197k versus about $204k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior AI Performance and Efficiency Engineer

As a Senior AI Performance and Efficiency Engineer, you will join the AI Efficiency team to enhance performance across the entire stack for researchers using GPU Clusters. You will collaborate with researchers to identify infrastructure and application deficiencies while building tools and frameworks to detect and analyze efficiency bottlenecks in diverse workloads including Robotics, Autonomous vehicles, LLMs, and Video. Your daily work involves monitoring fleet-wide utilization patterns and delivering scalable solutions to improve hardware, software, and infrastructure usage. To succeed, you must possess expertise in Python, Go, Bash, and cloud platforms like AWS, GCP, or Azure. You will utilize tools such as NSight Systems, NSight Compute, and NCCL for distributed training. Preferred skills include CUDA programming, MLPerf benchmarking, InfiniBand with RDMA, and experience with PyTorch, TensorFlow, Lustre, or GPFS storage systems.

What you'll do

  • Identify and resolve infrastructure and application deficiencies to improve ML model efficiency.
  • Build tools and frameworks to detect and analyze performance bottlenecks in GPU clusters.
  • Optimize training and inference performance across diverse workloads like LLMs, robotics, and autonomous vehicles.
  • Monitor fleet-wide utilization patterns to identify and solve large-scale inefficiency problems.
  • Develop scalable solutions to improve the usage of hardware, software, and infrastructure.
  • Debug and optimize distributed training using NCCL and tools like NSight Systems.
  • Advocate for and integrate the latest AI/ML technologies and frameworks into organizational workflows.

What we're looking for

  • BS or similar background in Computer Science or a related field (or equivalent experience).
  • Minimum 5+ years of experience designing and operating large scale compute infrastructure.
  • Strong understanding of modern ML techniques, tools, and deep learning concepts.
  • Experience investigating and resolving training and inference performance end to end.
  • Debugging and optimization experience with NSight Systems and NSight Compute.
  • Experience debugging large-scale distributed training using NCCL.
  • Proficiency in Python, Go, or Bash scripting and familiarity with cloud computing platforms like AWS, GCP, or Azure.
  • Familiarity with parallel computing frameworks, CUDA programming, and deep learning frameworks like PyTorch or TensorFlow.

More like this

Similar roles

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Senior AI Engineer

Intapp

Palo Alto, CA +2 149 days ago
LangGraph LangChain Python LLM CrewAI AutoGen Azure AWS Google Cloud Docker PostgreSQL CI/CD GitHub Actions Azure Pipelines
7+ yrs exp Hybrid

Senior AI Engineer

Abbott

Remote 93 days ago $99,300$198,700
Go AI SaaS RESTful APIs Microservices SQL Server PostgreSQL CI/CD Kubernetes Docker Terraform Linux MLOps Vector Databases Retrieval-Augmented Generation Distributed Systems NoSQL
8+ yrs exp Remote

Senior AI Engineer

Mastercard

O Fallon, MO 11 days ago $115,000$184,000
Generative AI LLM RAG LangGraph CrewAI AutoGen Python PyTorch TensorFlow Databricks MLflow Pinecone Neo4j AWS FastAPI SQL Hugging Face Prompt Engineering LoRA PEFT

Senior AI Engineer

Morgan Stanley

New York, NY 128 days ago $120,000$165,000
Python Generative AI LLM FastAPI Flask SQL Docker Kubernetes Kafka Redis Prometheus Grafana OpenTelemetry CI/CD GitOps Data Engineering multiprocessing multithreading Asynchronous I/O
5+ yrs exp

Senior AI Engineer

Allstate

Remote (IL) 120 days ago $100,000$170,500
Python LLMs Agentic AI RDF OWL SPARQL Knowledge Graphs Microsoft Fabric Azure Spark SQL ETL ELT Star Schema Microservices MLOps LLMOps CI/CD
6+ yrs exp Remote