Senior HPC and LSF Operations Engineer

Nvidia

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Santa Clara, CAWestford, MAAustin, TXDurham, NC
Salary
$152,000–$241,500 / yr
Posted
2 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $174k
This role $197k
$127k most similar roles pay here $254k

This role pays more than 71% of similar roles. Most pay $139,175–$209,725 — the shaded band above. At the midpoint, this role pays about $197k versus about $174k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior HPC and LSF Operations Engineer

As a Senior HPC and LSF Operations Engineer on the Hardware Infrastructure EDA Compute team, you will manage, scale, and optimize job scheduling systems in a multi-site environment. You will be responsible for analyzing performance data to identify bottlenecks, improving utilization and throughput, and resolving issues across scheduler, OS, and workload layers. The role involves implementing automation to reduce manual effort, defining reliable metrics and SLOs, and developing observability tools like monitoring pipelines and performance dashboards. You will work with technologies including LSF, Slurm, Linux systems administration (CentOS/RHEL), and container technologies such as Docker, Singularity, or Podman. This position focuses on the specific technical challenge of supporting compute-intensive workloads within an EDA environment to improve design velocity and infrastructure efficiency while ensuring high service reliability for complex system behaviors under load.

What you'll do

  • Manage, scale, and optimize job scheduling systems like LSF and Slurm in multi-site environments.
  • Analyze performance data to identify bottlenecks and improve system utilization, throughput, and turnaround time.
  • Resolve service-impacting issues across scheduler, operating system, and workload layers.
  • Implement automation and process improvements to address recurring operational challenges and reduce manual effort.
  • Define and track reliable metrics and SLOs for service performance and reliability.
  • Develop observability systems including monitoring pipelines, alerting strategies, and performance dashboards.
  • Create technical documentation and establish best practices to ensure consistency across multiple sites.
  • Partner with customer teams to clarify requirements and communicate technical tradeoffs regarding infrastructure.

What we're looking for

  • Bachelor’s degree in Computer Science or a related field, or equivalent experience.
  • Minimum 5+ years of experience operating and supporting large-scale Linux-based compute infrastructure.
  • Strong hands-on experience supporting and tuning job scheduling systems like LSF or Slurm in HPC or silicon design environments.
  • Proficiency in Linux systems administration using CentOS/RHEL.
  • Ability to independently analyze complex system behavior under load and solve problems across scheduler, OS, and workload layers.
  • Clear and effective communication skills to articulate technical tradeoffs and reliability metrics to stakeholders.
  • Experience implementing reliability engineering practices or building observability systems for HPC environments.
  • Familiarity with container technologies such as Docker, Singularity, or Podman in HPC environments.

More like this

Similar roles

HPC Operations Engineer

Nvidia

Santa Clara, CA +3 4 days ago $124,000$195,500
Linux RHEL CentOS Ubuntu Python Bash Slurm LSF NFS LDAP HPC EDA Automation Scripting
2+ yrs exp Hybrid

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

HPC Engineer

General Dynamics

Rockville, MD 29 days ago $123,250$166,750
HPC Linux Slurm Python Bash Spack EasyBuild Apptainer Singularity InfiniBand Scientific Software Module Systems Compilers Package Management Networking
8+ yrs exp Hybrid