Senior Solutions Architect, AI Factory Observability and Visualization

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Austin, TXDurham, NCSanta Clara, CA
Salary
$184,000–$287,500 / yr
Posted
79 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $207k
This role $236k
$154k most similar roles pay here $302k

This role pays more than 80% of similar roles. Most pay $178,700–$235,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $207k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Solutions Architect, AI Factory Observability and Visualization

Senior Solutions Architect, AI Factory Observability and Visualization joins the Infrastructure Specialists team to develop full-spectrum visibility for HPC systems and AI factories. This role involves running validation tools, microbenchmarks, and workloads to assess system health while establishing metrics, logs, and signals across the stack. The architect will build and extend the telemetry surface for hardware, fabric, and workload components, ensuring data is collected, transformed, and surfaced effectively. Key responsibilities include developing automation using Python and Shell to manage network and system data. Technical requirements include experience with Linux-based systems, multi-GPU clusters, and observability tools like Prometheus, Grafana, and Loki. The role focuses on the technical challenge of transforming complex telemetry from GPU and fabric components into actionable insights for distributed environments, ensuring infrastructure readiness through collaboration with hardware, software, and networking groups.

What does a Solutions Architect earn in Texas?

Median $182825 from 30 postings across 11 companies.

See salary data

What you'll do

  • Run validation tools, microbenchmarks, and workloads to assess system health and performance.
  • Define "healthy" system states by identifying key metrics, logs, and signals across the stack.
  • Build and extend telemetry surfaces for hardware, fabric, and workload data collection.
  • Develop Python and Shell scripts to automate the collection and transformation of system data.
  • Investigate visibility gaps to ensure observability tools accurately reflect true system behavior.
  • Transform complex metrics, logs, and traces into actionable insights for distributed environments.
  • Recommend improvements to data sources and reporting to provide clearer insight into system performance.

What we're looking for

  • Bachelor's degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or a related field.
  • 6+ years of experience managing Linux-based systems in HPC, distributed systems, or large AI/ML settings.
  • Hands-on experience with the architecture of multi-GPU and/or multi-node clusters, including networking and interconnects.
  • Proficiency in Python and Shell/Bash for scripting, automation, and tooling.
  • Practical experience with observability systems such as Prometheus, Grafana, or Loki, including building custom exporters and handling metric cardinality.
  • Experience transforming metrics, logs, and traces into actionable insights for complex distributed environments.
  • Familiarity with GPU and fabric telemetry, such as DCGM, NVLink, and InfiniBand/Ethernet fabric counters.
  • Strong communication skills to collaborate effectively with cross-functional hardware, software, networking, and product teams.

More like this

Similar roles

Senior Solutions Architect, AI Infrastructure

Nvidia

Remote (Santa Clara, CA) 71 days ago $184,000$287,500
GPU NVLink HPC Distributed Systems Networking NCCL MPI IMEX NMX Cluster Design Performance Modeling AI Infrastructure
8+ yrs exp Remote

Senior Solutions Architect, AI Infrastructure

Nvidia

Remote (Santa Clara, CA) 68 days ago $184,000$287,500
GPU NVLink HPC Distributed Systems Networking NCCL MPI IMEX NMX Cluster Design Performance Modeling AI Infrastructure
8+ yrs exp Remote