Senior Solutions Architect, AI Factory Deployment

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Austin, TXDurham, NCSanta Clara, CA
Salary
$124,000–$195,500 / yr
Posted
7 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $199k
This role $160k
$110k most similar roles pay here $256k

This role pays less than 78% of similar roles. Most pay $162,000–$235,750 — the shaded band above. At the midpoint, this role pays about $160k versus about $199k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Solutions Architect, AI Factory Deployment

As a Senior Solutions Architect, AI Factory Deployment - NVIS, you will join the NVIDIA Infrastructure Specialists team to support the creation, implementation, and verification of AI factories. You will manage multi-GPU and multi-node Linux clusters, focusing on running and debugging AI/LLM workloads and benchmarks. Your daily responsibilities include validating configurations for NCCL and collectives like AllReduce and AllToAll, investigating failed training jobs, and building observability tools such as metrics, logs, and dashboards. To improve performance and scalability, you will develop automation using Python and Shell to conduct benchmarks and regression checks. You will collaborate with cross-functional hardware, software, and networking teams to prepare infrastructure for customers. The role requires expertise in distributed systems, HPC performance engineering, and analyzing communication patterns to optimize throughput, latency, and scaling efficiency for large-scale AI applications.

What does a Solutions Architect earn in Texas?

Median $182825 from 30 postings across 11 companies.

See salary data

What you'll do

  • Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
  • Validate configurations for NCCL, collective communications (AllReduce/AllToAll), and distributed training frameworks.
  • Execute, orchestrate, and analyze key AI and LLM benchmarks to evaluate system performance.
  • Troubleshoot and resolve issues when training jobs or benchmarks fail, hang, or perform poorly.
  • Develop automation scripts using Python and Shell for benchmarking, result collection, and regression checks.
  • Build observability tools including metrics, logs, and dashboards to monitor workload behavior and system health.
  • Analyze communication patterns to recommend improvements for job configuration, parallelism strategies, and scaling efficiency.
  • Collaborate with cross-functional teams to prepare AI factories and documentation for customer use.

What we're looking for

  • Bachelor's degree in Computer Science, Mathematics, Engineering, Physics, or a related field.
  • 5+ years of experience managing Linux-based systems in HPC, distributed systems, or AI/ML environments.
  • Hands-on experience running AI/ML workloads on multi-GPU and multi-node clusters.
  • Experience with NCCL and collective communication patterns like AllReduce and AllToAll.
  • Proficiency in Python and Shell/Bash for scripting, automation, and tooling.
  • Experience benchmarking distributed systems and performing performance analysis for GPU-accelerated environments.
  • Familiarity with observability stacks including metrics, logging, and tracing for large distributed systems.
  • Strong communication skills to collaborate effectively with cross-functional teams.

More like this

Similar roles

Senior Solutions Architect, AI Infrastructure

Nvidia

Remote (Santa Clara, CA) 71 days ago $184,000$287,500
GPU NVLink HPC Distributed Systems Networking NCCL MPI IMEX NMX Cluster Design Performance Modeling AI Infrastructure
8+ yrs exp Remote

Senior Solutions Architect, AI Infrastructure

Nvidia

Remote (Santa Clara, CA) 68 days ago $184,000$287,500
GPU NVLink HPC Distributed Systems Networking NCCL MPI IMEX NMX Cluster Design Performance Modeling AI Infrastructure
8+ yrs exp Remote