HPC Infrastructure & Cluster Engineer

General Dynamics

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Springfield, VA
Salary
$119,850–$162,150 / yr
Posted
9 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $186k
This role $141k
$106k most similar roles pay here $248k

This role pays less than 87% of similar roles. Most pay $152,950–$218,250 — the shaded band above. At the midpoint, this role pays about $141k versus about $186k for comparable roles.

Based on 240 similar postings.

Employer

About General Dynamics

General Dynamics is a global aerospace and defense company offering a broad portfolio of products and services in business aviation, ship construction, land combat vehicles, and information technology. It serves customers in the U.S. government, allied governments, and a diverse array of commercial markets.

General Dynamics currently has 1123 open roles on FindRole.

Listed pay typically runs $124,093–$150,385 across 984 roles with salary data.

Most-posted roles

View all roles at General Dynamics

At a glance

TL;DR · HPC Infrastructure & Cluster Engineer

As an HPC Infrastructure & Cluster Engineer within the User Facing and Data Center Services team, you will manage the administration, health, and performance of a dedicated customer compute cluster. You will perform end-to-end administration of the hardware foundation, including Linux operating system management, patching, and system upgrades to ensure a highly available environment for complex workloads. Your daily work involves optimizing infrastructure at the hardware, OS, and network levels, managing storage solutions, and maintaining high-speed InfiniBand GPU-to-GPU networks. You will utilize tools such as Run:AI, SLURM, and Red Hat OpenShift to manage job scheduling and container platforms. Technical requirements include proficiency in Bash and Python for automation, experience with bare-metal servers, and expertise in managing parallel file systems to solve performance engineering challenges for intensive AI/ML workloads.

What you'll do

  • Manage daily operations of the compute cluster including Linux administration, hardware monitoring, patching, and system upgrades.
  • Configure and optimize workload management platforms using Run:AI to distribute AI/ML workloads across the cluster.
  • Tune infrastructure performance at the hardware, operating system, and network levels to maximize data throughput.
  • Administer storage solutions and manage high-speed InfiniBand GPU-to-GPU networking infrastructure.
  • Provision environment dependencies and container platforms specifically using Red Hat OpenShift for customer model deployment.
  • Ensure all infrastructure components comply with federal security standards and maintain necessary system accreditations.
  • Develop automation and configuration scripts in Bash or Python to streamline cluster maintenance tasks.

What we're looking for

  • Must be a United States citizen.
  • Must possess an active Top Secret/SCI clearance with the ability to obtain a CI Polygraph.
  • Must have 5+ years of experience in Linux systems administration and infrastructure management for high-performance computing environments.
  • Must have expertise managing bare-metal servers, enterprise storage arrays, and advanced network configurations including InfiniBand.
  • Must be proficient with workload managers, job schedulers, and AI orchestration tools such as Run:AI or SLURM.
  • Must have hands-on experience with enterprise container orchestration platforms like OpenShift or Kubernetes.
  • Must be able to write automation and configuration scripts using Bash or Python.
  • US Citizenship.

More like this

Similar roles

Senior HPC AI Cluster Engineer

Nvidia

Remote 22 days ago $176,000$276,000
HPC AI GPU CUDA Slurm Kubernetes Python Bash Ansible Jenkins InfiniBand Ethernet RDMA Lustre GPFS Weka.io Linux RedHat CentOS Ubuntu AWS Azure Google Cloud VMware KVM
8+ yrs exp Remote

HPC Engineer

General Dynamics

Rockville, MD 29 days ago $123,250$166,750
HPC Linux Slurm Python Bash Spack EasyBuild Apptainer Singularity InfiniBand Scientific Software Module Systems Compilers Package Management Networking
8+ yrs exp Hybrid

Software Engineer, Compute Infra / HPC

Microsoft

37 days ago $142,800$274,800
Kubernetes Go Rust C++ C# Python Terraform Infrastructure as Code Distributed Systems Linux Azure InfiniBand RoCE RDMA NVLink NCCL Bare-metal Provisioning Container Runtimes
4+ yrs exp Hybrid

AI/HPC Cluster Design Engineer

Amd

Austin, TX 53 days ago $133,200$199,800
HPC AI Systems GPU CPU InfiniBand Ethernet Lustre Ceph PCIe UALink Cluster Design Data Center Engineering Power Delivery System Architecture

HPC Performance Engineer

Nvidia

Remote (OR) +3 53 days ago $152,000$241,500
HPC CUDA C++ C Fortran OpenMP MPI OpenACC Compiler Optimization Assembly Linear Algebra Numerical Methods Multi-GPU Systems
5+ yrs exp Remote

HPC Systems Engineer, Modeling & Simulation

Anduril Industries

Costa Mesa, CA 92 days ago $132,000$198,000
HPC Linux Unix Python Bash Slurm MPI OpenMP CMake NFS NAS TCP/IP ParaView VisIt Cubit CTH ALE3D Sierra Data Workflows
5+ yrs exp