Senior Solutions Architect, AI Cluster Performance and Telemetry

Nvidia

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Santa Clara, CAAustin, TX
Salary
$184,000–$287,500 / yr
Posted
99 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $206k
This role $236k
$159k most similar roles pay here $301k

This role pays more than 74% of similar roles. Most pay $172,413–$239,606 — the shaded band above. At the midpoint, this role pays about $236k versus about $206k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Solutions Architect, AI Cluster Performance and Telemetry

As a Senior Solutions Architect, AI Cluster Performance and Telemetry, you will join the solutions architecture team to serve as a technical expert at the intersection of hardware and complex software stacks. You will analyze and optimize performance for AI, deep learning, and HPC ecosystems by identifying bottlenecks across interconnected GPU, CPU, and networking systems. Your daily responsibilities include maintaining benchmarking suites, monitoring hardware counters, and investigating system configurations to resolve discrepancies affecting peak performance. You will utilize tools such as Perf, eBPF, Prometheus, Grafana, Docker, Kubernetes, SLURM, and Ansible while working with NCCL and NVIDIA architectures like NVLink. The role focuses on the technical challenge of managing high-performance clusters, transforming raw telemetry into structured data, and optimizing distributed AI training workloads and large language models within massive infrastructure environments.

What does a Solutions Architect earn in California?

Median $235750 from 45 postings across 8 companies.

See salary data

What you'll do

  • Identify and resolve complex performance bottlenecks across interconnected GPU, CPU, and networking systems.
  • Develop and maintain robust benchmarking suites to stress-test high-performance clusters and establish baselines.
  • Use industry-standard tools to monitor hardware performance counters and extract deep system telemetry.
  • Investigate system and software configurations to identify and fix discrepancies impacting peak performance.
  • Transform raw logs and telemetry into structured time series data, dashboards, and heat maps.
  • Translate complex technical performance anomalies into clear, actionable narratives for cross-functional teams.
  • Optimize distributed AI training workloads, LLMs, and large-scale high-performance computing environments.

What we're looking for

  • BS or MS in Engineering, Electrical Engineering, Physics, or Computer Science (or equivalent experience).
  • 8+ years of work-related experience in the high-tech industry, specifically in system build, performance analysis, and technical customer-facing roles.
  • Strong understanding of how CPUs, GPUs, and high-speed networking fabrics interact within massive clusters.
  • Practical experience with performance counters, profiling tools, and telemetry collection systems like Perf, eBPF, Prometheus, and Grafana.
  • Practical experience with containers, cloud provisioning, and scheduling tools such as Docker, Kubernetes, SLURM, and Ansible.
  • Proven track record of transforming raw logs and telemetry into structured time series data, dashboards, and heat maps.
  • Ability to translate complex technical performance anomalies into clear narratives for cross-functional teams.
  • Experience with multi-GPU communication libraries (NCCL), NVIDIA hardware architectures, or optimizing distributed AI training workloads.

More like this

Similar roles

Senior Solutions Architect, AI Infrastructure

Nvidia

Remote (Santa Clara, CA) 71 days ago $184,000$287,500
GPU NVLink HPC Distributed Systems Networking NCCL MPI IMEX NMX Cluster Design Performance Modeling AI Infrastructure
8+ yrs exp Remote

Senior Solutions Architect, AI Infrastructure

Nvidia

Remote (Santa Clara, CA) 68 days ago $184,000$287,500
GPU NVLink HPC Distributed Systems Networking NCCL MPI IMEX NMX Cluster Design Performance Modeling AI Infrastructure
8+ yrs exp Remote

Senior Solutions Architect, Supercomputing

Nvidia

Remote (TX) +1 157 days ago $184,000$287,500
GPU CUDA Machine Learning Deep Learning High-Performance Computing Generative AI Agentic AI Docker Kubernetes Slurm GPGPU Parallel Computing Data Science Networking
8+ yrs exp Remote

Senior Solutions Architect, Supercomputing

Nvidia

Remote (TX) +1 157 days ago $184,000$287,500
GPU CUDA Machine Learning Deep Learning High-Performance Computing Generative AI Agentic AI Docker Kubernetes Slurm GPGPU Data Science Parallel Computing Networking
8+ yrs exp Remote