Senior Product Architect, K8s-Based AI Infrastructure

Nvidia

Confirmed live 2 days ago High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$184,000–$287,500 / yr
Posted
22 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $206k
This role $236k
$154k most similar roles pay here $302k

This role pays more than 73% of similar roles. Most pay $169,200–$242,212 — the shaded band above. At the midpoint, this role pays about $236k versus about $206k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Product Architect, K8s-Based AI Infrastructure

As a Senior Product Architect, K8s-Based AI Infrastructure, you will join the team to design and shape architectures connecting powerful AI clusters. You will transform ideas into functional products by defining detailed architectures for performance, scalability, interoperability, and datacenter deployment. Your daily responsibilities include leading prototyping and testing to diagnose bottlenecks, iterating on designs based on feedback, and collaborating with hardware, software, security, and systems teams to integrate networking, compute, and storage ecosystems. You will also build proof-of-concepts and minimum viable products while creating technical documentation and whitepapers. The role requires expertise in datacenter-scale HPC or AI infrastructure, Python, Ansible, Container Runtimes, and Kubernetes. You will address complex problems involving agentic and RAG-based workflows, inference at scale, large-scale training, fine-tuning, and model evaluation within the domain of AI infrastructure and Agentic AI.

What you'll do

  • Propose new product concepts based on emerging trends in AI Infrastructure and Agentic AI.
  • Define detailed product architectures focusing on performance, scalability, interoperability, and datacenter deployment.
  • Lead prototyping and testing processes to validate design functionality and resolve performance bottlenecks.
  • Iterate and refine designs based on technical feedback, test results, and evolving requirements.
  • Coordinate with internal teams to ensure seamless integration of networking, compute, and storage ecosystems.
  • Develop proof-of-concepts and minimum viable products in collaboration with partners and customers.
  • Create comprehensive technical documentation including specifications, user guides, whitepapers, and educational content.

What we're looking for

  • Bachelor's degree in computer science or a related field (or equivalent experience).
  • 8+ years of experience architecting datacenter-scale HPC or AI infrastructure as a Principal Architect, Solutions Architect, Principal Engineer, or equivalent.
  • Strong background in technologies for large scale management of complex systems, networks, and storage.
  • Extensive experience with DevOps solutions including Python, Ansible, Container Runtimes, Kubernetes, and data center deployments.
  • Familiarity with AI workloads, including agentic & RAG-based workflows, inference at scale, training, fine-tuning, and model evaluation.
  • Exceptional communication skills to translate complex technical details for diverse audiences.
  • Prior first-hand experience building large scale AI infrastructure (preferred).
  • Advanced certifications or publications in AI, deep learning, or related fields (preferred).

More like this

Similar roles

Senior Solutions Architect, Agentic AI

Nvidia

Remote 51 days ago $184,000$287,500
Agentic AI LLM RAG Python C/C++ PyTorch LangChain LlamaIndex CrewAI LangGraph NVIDIA NIM NeMo Framework TensorRT-LLM Triton Inference Server CUDA Kubernetes CI/CD AWS GCP Azure Spark Dask
8+ yrs exp Remote

AI Factory CPU focused Solutions Architect

Nvidia

Remote 58 days ago $184,000$287,500
HPC AI Infrastructure MLOps Arm CPU Networking Reference Architectures Agentic AI Reinforcement Learning Inference Workflows Automation Performance Testing Data Center Architecture Proof-of-Concept
Remote

Senior Solutions Architect, AI Infrastructure

Nvidia

Remote (Santa Clara, CA) 71 days ago $184,000$287,500
GPU NVLink HPC Distributed Systems Networking NCCL MPI IMEX NMX Cluster Design Performance Modeling AI Infrastructure
8+ yrs exp Remote