HPC Operations Engineer

Nvidia

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CAWestford, MAAustin, TXDurham, NC
Salary
$124,000–$195,500 / yr
Posted
4 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $178k
This role $160k
$113k most similar roles pay here $230k

This role pays less than 65% of similar roles. Most pay $142,400–$214,000 — the shaded band above. At the midpoint, this role pays about $160k versus about $178k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · HPC Operations Engineer

As an HPC Operations Engineer, you will join a team dedicated to maintaining the high-performance computing environment that supports semiconductor build and advanced engineering workflows. Your daily responsibilities include providing first-line support for users regarding scheduling, compute, storage, and access issues while troubleshooting job failures, resource constraints, and performance concerns. You will monitor system health, perform infrastructure triage, execute maintenance procedures like patching and configuration updates, and develop technical documentation and runbooks. The role requires proficiency in Linux systems administration across RHEL, CentOS, and Ubuntu platforms. Candidates should possess skills in Bash or Python scripting, experience with workload schedulers such as LSF or Slurm, and knowledge of network computing infrastructure including NFS, automounter, and LDAP. You will specifically address technical challenges within the domain of EDA workloads and large-scale compute environments.

What you'll do

  • Provide first-line support for HPC users regarding scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, resource constraints, and performance concerns to resolution.
  • Triage infrastructure incidents by gathering diagnostics and escalating complex issues to subject matter experts.
  • Monitor system health, queue status, node availability, and service stability for daily operations.
  • Execute operational procedures for system maintenance, patching, and configuration updates.
  • Create and maintain technical documentation, runbooks, and knowledge base articles for users and internal teams.
  • Identify recurring issues to propose and implement practical workflow improvements.

What we're looking for

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent experience.
  • 2+ years of experience supporting Linux-based production environments.
  • Solid Linux systems administration fundamentals for RHEL, CentOS, and/or Ubuntu.
  • Experience interacting directly with users in a technical support or operations role.
  • Strong written communication skills to produce clear documentation and procedural guides.
  • Foundational scripting or automation experience using Bash or Python.
  • Solid understanding of workload schedulers such as LSF, Slurm, or similar systems.
  • Strong grasp of network computing supporting infrastructure including NFS, automounter, and LDAP.

More like this

Similar roles

Senior HPC and LSF Operations Engineer

Nvidia

Santa Clara, CA +3 2 days ago $152,000$241,500
LSF Slurm Linux HPC CentOS RHEL Docker Singularity Podman Observability Monitoring Automation Metrics
5+ yrs exp Hybrid

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

HPC Engineer

General Dynamics

Rockville, MD 29 days ago $123,250$166,750
HPC Linux Slurm Python Bash Spack EasyBuild Apptainer Singularity InfiniBand Scientific Software Module Systems Compilers Package Management Networking
8+ yrs exp Hybrid

HPC Performance Engineer

Nvidia

Remote (OR) +3 53 days ago $152,000$241,500
HPC CUDA C++ C Fortran OpenMP MPI OpenACC Compiler Optimization Assembly Linear Algebra Numerical Methods Multi-GPU Systems
5+ yrs exp Remote

HPC Systems Engineer, Modeling & Simulation

Anduril Industries

Costa Mesa, CA 92 days ago $132,000$198,000
HPC Linux Unix Python Bash Slurm MPI OpenMP CMake NFS NAS TCP/IP ParaView VisIt Cubit CTH ALE3D Sierra Data Workflows
5+ yrs exp