Solutions Architect, Inference Deployments

Nvidia

Confirmed live 3 days ago High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$152,000–$241,500 / yr
Posted
150 days ago
Freshness
Confirmed live 3 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $199k
This role $197k
$134k most similar roles pay here $253k

This role pays more than 58% of similar roles. Most pay $162,000–$235,750 — the shaded band above. At the midpoint, this role pays about $197k versus about $199k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Solutions Architect, Inference Deployments

As a Solutions Architect, Inference Deployments, you will join a team focused on rolling out and enhancing AI inference solutions at scale using GPU technology and Kubernetes. You will collaborate with engineering, DevOps teams, and customers to build inference pipelines, distribute tasks among GPU workers for efficiency, and orchestrate disaggregated inference for complex workloads. Your daily work involves accelerating pipelines using TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo while providing technical leadership and mentorship to resolve complex deployment issues. You will utilize tools like Triton Inference Server, NVIDIA GPU Operator, NIM Operator, and Multi-Instance GPU partitioning. The role requires expertise in solving GPU allocation, memory hierarchies, and low-latency networking using RDMA and UCX. You will focus on the technical challenge of tuning large language models for low-latency inference within enterprise environments.

What does a Solutions Architect earn in California?

Median $235750 from 45 postings across 8 companies.

See salary data

What you'll do

  • Build inference pipelines using tools like NVIDIA Dynamo to distribute tasks among GPU workers for improved efficiency.
  • Orchestrate disaggregated inference workloads using Kubernetes in collaboration with DevOps teams.
  • Accelerate inference pipelines using TensorRT-LLM, vLLM, and SGLang to ensure seamless integration.
  • Provide technical leadership and mentorship to customers during the deployment of disaggregated inference systems.
  • Resolve complex technical issues related to GPU allocation, memory hierarchies, and low-latency networking.
  • Tune large language models for low-latency inference in enterprise environments.
  • Optimize model serving using technologies like Triton Inference Server or NVIDIA NIM.

What we're looking for

  • At least 5 years of experience in Solutions Architecture with a track record of deploying distributed systems and AI inference workloads on Kubernetes.
  • Experience using NVIDIA Dynamo, Triton Inference Server, or TensorRT-LLM for model optimization and serving.
  • Expertise in GPU orchestration using NVIDIA GPU Operator, NIM Operator, and Multi-Instance GPU (MIG) partitioning.
  • Proficiency in solving complex GPU allocation, memory hierarchies, and low-latency networking such as RDMA and UCX.
  • Proven success tuning large language models for low-latency inference in enterprise environments.
  • A Bachelor of Science in Computer Science, Engineering, or equivalent experience.
  • Knowledge of transformer neural networks and acceleration technologies like quantization and speculative decoding is preferred.
  • NVIDIA Certified AI Engineer or similar professional credentials are preferred.

More like this

Similar roles

Solutions Architect, Inference Deployments

Nvidia

Santa Clara, CA 150 days ago $152,000$241,500
Kubernetes NVIDIA GPU TensorRT-LLM vLLM SGLang Triton Inference Server NVIDIA Dynamo KServe RDMA UCX Quantization Speculative Decoding Multi-Instance GPU (MIG)
5+ yrs exp

Senior Solutions Architect, Generative AI Deployment and AIOps

Nvidia

Remote 15 days ago $184,000$287,500
Generative AI LLMs Deep Learning PyTorch TensorFlow Python C++ Kubernetes MLOps NVIDIA NIM TensorRT TensorRT-LLM GPU Orchestration Multi-Instance GPU (MIG) Distributed Computing Performance Analysis Profiling
Remote

Solutions Architect, AI Models

Nvidia

Remote (Santa Clara, CA) 19 days ago $152,000$241,500
PyTorch JAX TensorFlow Hugging Face Transformers Python Linux Kubernetes SLURM NeMo Nemotron Distributed Computing LLMs Reinforcement Learning Model Optimization Data Processing
Remote

Senior Solutions Architect, Physical AI Cloud

Nvidia

Remote (Santa Clara, CA) 33 days ago $152,000$241,500
Kubernetes TensorRT-LLM vLLM SGLang Triton Isaac Sim Isaac Lab ROS2 Airflow Argo GitOps IaC REST gRPC S3 NFS Lustre
5+ yrs exp Remote