Solutions Architect, Inference Deployments
Nvidia
Quick summary
Market check
How this pay compares to similar roles
This role pays more than 58% of similar roles. Most pay $162,000–$235,750 — the shaded band above. At the midpoint, this role pays about $197k versus about $199k for comparable roles.
Based on 239 similar postings.
Employer
Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing
Nvidia currently has 896 open roles on FindRole.
Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.
Most-posted roles
At a glance
As a Solutions Architect, Inference Deployments, you will join a team focused on rolling out and enhancing AI inference solutions at scale using GPU technology and Kubernetes. You will collaborate with engineering, DevOps teams, and customers to build inference pipelines, distribute tasks among GPU workers, and orchestrate disaggregated inference for complex workloads. Your daily work involves accelerating inference pipelines using TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo while providing technical leadership and mentorship to resolve complex deployment issues. You will utilize tools like Triton Inference Server, NVIDIA GPU Operator, NIM Operator, and Multi-Instance GPU partitioning. The role requires expertise in solving sophisticated GPU allocation, memory hierarchies, and low-latency networking using RDMA and UCX. You will focus on the technical challenge of tuning large language models for low-latency inference within enterprise environments to ensure seamless production integration.
What does a Solutions Architect earn in California?
Median $235750 from 45 postings across 8 companies.
Skills
What you'll do
What we're looking for
More like this
Nvidia
Nvidia
Nvidia
Nvidia
Nvidia
Nvidia