Senior HPC Support Engineer, InfiniBand - NVLink

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Westford, MADurham, NCRedmond, WASanta Clara, CANew York, NY
Salary
$108,000–$172,500 / yr
Posted
37 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $194k
This role $140k
$93k most similar roles pay here $251k

This role pays less than 84% of similar roles. Most pay $155,675–$231,437 — the shaded band above. At the midpoint, this role pays about $140k versus about $194k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior HPC Support Engineer, InfiniBand - NVLink

As a Senior HPC Support Engineer, InfiniBand - NVLink, you will join the NVIDIA Experience Global Technical Support team to provide comprehensive solutions for sophisticated installations, maintenance, and operations of groundbreaking networking products. You will serve as a primary point of contact for customers, resolving technical issues through meticulous research and reproduction while collaborating with Engineering and Marketing teams on product requirements. The role involves troubleshooting systems using multi-distribution Linux operating systems with a focus on InfiniBand, NVLink, and GPU technology. Required expertise includes networking protocols like IP, L2, and L3; tools such as TCPDUMP and Wireshark; and experience with RDMA/RoCEv2, NCCL, MPI, and Slurm. You will also utilize Agentic AI technologies like Claude or Cursor while managing complex infrastructure involving containerized solutions, virtualization, and cloud platforms to solve high-level networking and AI infrastructure problems.

What you'll do

  • Resolve complex customer technical issues regarding InfiniBand, NVLink, and GPU technologies.
  • Provide primary point of contact for customers via telephone, email, and conference calls.
  • Debug networking protocols using tools like TCPDUMP and Wireshark to resolve installation and operation problems.
  • Troubleshoot multi-vendor interoperability and performance issues in large-scale AI infrastructure environments.
  • Develop and document standard methodologies for internal support and R&D teams to improve processes.
  • Provide feedback to engineering and marketing teams regarding product requirements and customer experience.
  • Manage technical support for Linux-based systems across multiple distributions.

What we're looking for

  • Must have 5+ years of experience in customer support and debugging for large-scale networking and AI infrastructure environments.
  • Must possess an academic degree in Networking, Computer Science/Engineering, Electrical Engineering, or IT (or equivalent experience).
  • Must demonstrate proficiency with InfiniBand, NVLink, Ethernet, and GPU technologies.
  • Must have expertise in Linux OS system administration, networking, and performance at a LFCS/RHCSA level.
  • Must be able to debug networking protocols using tools such as TCPDUMP and Wireshark.
  • Must possess knowledge of containerized solutions (DCA/CKA), virtualization (KVM/ESXi), and cloud infrastructure (AWS/OCI).
  • Must have experience with high-performance computing technologies including RDMA, RoCEv2, NCCL, MPI, and Slurm.
  • Must be proficient in shell scripting using Bash or Python.

More like this

Similar roles

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

HPC Operations Engineer

Nvidia

Santa Clara, CA +3 4 days ago $124,000$195,500
Linux RHEL CentOS Ubuntu Python Bash Slurm LSF NFS LDAP HPC EDA Automation Scripting
2+ yrs exp Hybrid

Senior Solution Engineer, Networking

Nvidia

Santa Clara, CA +4 35 days ago $168,000$270,250
C Python InfiniBand NVLink Spectrum-X Linux DPDK SRIOV SDN CUDA NCCL MPI DOCA Firmware NIC Drivers Switch ASICs Embedded Systems HPC Networking

Senior Solutions Architect, AI Compute

Nvidia

Remote 21 days ago $184,000$287,500
Linux Python Bash Ansible Kubernetes SLURM LSF UGE InfiniBand MPI HPL NCCL MLPerf Lustre GPFS GPU HPC Networking System Administration Automation
8+ yrs exp Remote