Solutions Architect, Inference Deployments

Nvidia

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$152,000–$241,500 / yr
Posted
149 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $199k
This role $197k
$134k most similar roles pay here $253k

This role pays more than 58% of similar roles. Most pay $162,000–$235,750 — the shaded band above. At the midpoint, this role pays about $197k versus about $199k for comparable roles.

Based on 239 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Solutions Architect, Inference Deployments

As a Solutions Architect, Inference Deployments, you will join a team focused on rolling out and enhancing AI inference solutions at scale using GPU technology and Kubernetes. You will collaborate with engineering, DevOps teams, and customers to build inference pipelines, distribute tasks among GPU workers, and orchestrate disaggregated inference for complex workloads. Your daily work involves accelerating inference pipelines using TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo while providing technical leadership and mentorship to resolve complex deployment issues. You will utilize tools like Triton Inference Server, NVIDIA GPU Operator, NIM Operator, and Multi-Instance GPU partitioning. The role requires expertise in solving sophisticated GPU allocation, memory hierarchies, and low-latency networking using RDMA and UCX. You will focus on the technical challenge of tuning large language models for low-latency inference within enterprise environments to ensure seamless production integration.

What does a Solutions Architect earn in California?

Median $235750 from 45 postings across 8 companies.

See salary data

What you'll do

  • Build inference pipelines using tools like NVIDIA Dynamo to distribute tasks among GPU workers for improved efficiency.
  • Orchestrate disaggregated inference workloads using Kubernetes in collaboration with DevOps teams.
  • Accelerate inference pipelines using TensorRT-LLM, vLLM, SGLang, and other backends.
  • Provide technical leadership and mentorship to customers during the deployment of disaggregated inference systems.
  • Resolve complex technical issues related to GPU allocation, memory hierarchies, and low-latency networking.
  • Tune large language models for low-latency inference in enterprise environments.
  • Implement model optimization techniques such as quantization and speculative decoding.

What we're looking for

  • At least 5 years of experience in Solutions Architecture with a track record of deploying distributed systems and AI inference workloads on Kubernetes.
  • Experience with NVIDIA Dynamo, Triton Inference Server, or TensorRT-LLM for model optimization and serving.
  • Expertise in GPU orchestration using NVIDIA GPU Operator, NIM Operator, and Multi-Instance GPU (MIG) partitioning.
  • Proficiency in solving complex GPU allocation, memory hierarchies, and low-latency networking like RDMA and UCX.
  • Demonstrated success tuning large language models for low-latency inference in enterprise environments.
  • A Bachelor of Science in Computer Science, Engineering, or equivalent experience.
  • Knowledge of transformer neural networks and acceleration technologies such as quantization and speculative decoding.
  • NVIDIA Certified AI Engineer or similar professional credentials.

More like this

Similar roles

Solutions Architect, Inference Deployments

Nvidia

Santa Clara, CA 149 days ago $152,000$241,500
Kubernetes NVIDIA GPU TensorRT-LLM vLLM SGLang Triton Inference Server NVIDIA Dynamo KServe RDMA UCX Quantization Speculative Decoding Multi-Instance GPU (MIG)
5+ yrs exp

Senior Solutions Architect, Generative AI Deployment and AIOps

Nvidia

Remote 14 days ago $184,000$287,500
Generative AI LLMs Deep Learning PyTorch TensorFlow Python C++ Kubernetes MLOps NVIDIA NIM TensorRT TensorRT-LLM GPU Orchestration Multi-Instance GPU (MIG) Distributed Computing Performance Analysis Profiling
Remote

Solutions Architect, AI Models

Nvidia

Remote (Santa Clara, CA) 18 days ago $152,000$241,500
PyTorch JAX TensorFlow Hugging Face Transformers Python Linux Kubernetes SLURM NeMo Nemotron Distributed Computing LLMs Reinforcement Learning Model Optimization Data Processing
Remote

Senior Solutions Architect, Physical AI Cloud

Nvidia

Remote (Santa Clara, CA) 32 days ago $152,000$241,500
Kubernetes TensorRT-LLM vLLM SGLang Triton Isaac Sim Isaac Lab ROS2 Airflow Argo GitOps IaC REST gRPC S3 NFS Lustre
5+ yrs exp Remote