NCX Senior Engineer

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CASeattle, WA
Salary
$184,000–$287,500 / yr
Posted
19 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $180k
This role $236k
$120k most similar roles pay here $305k

This role pays more than 90% of similar roles. Most pay $146,395–$213,375 — the shaded band above. At the midpoint, this role pays about $236k versus about $180k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · NCX Senior Engineer

The NCX Senior Engineer joins the DSX team to manage and improve operational capabilities for large-scale NVIDIA accelerated infrastructure in production. This role focuses on guiding NVIDIA Cloud Partners through Day 2 operations, including infrastructure health, observability, lifecycle management, and automated remediation. You will build continuous validation systems for GPU, CPU, storage, and network components while developing telemetry, monitoring, and alerting across Kubernetes and InfiniBand/RoCE networking environments. Key responsibilities include managing firmware lifecycles, OS patching, and translating reference architectures into production runbooks. The ideal candidate possesses expertise in Linux-based distributed systems, cloud infrastructure, and site reliability engineering. Required technical skills include proficiency in Python, Go, or shell scripting to automate infrastructure management. This role addresses the complex challenge of maintaining reliable performance for large-scale AI training and inference workloads across diverse partner environments.

What you'll do

  • Lead Day 2 operational readiness efforts for NVIDIA Cloud Partners to manage large-scale accelerated infrastructure.
  • Develop and implement methods to continuously validate GPU, CPU, storage, and network health across AI clusters.
  • Establish comprehensive observability, monitoring, alerting, and dashboards for compute, networking, and Kubernetes workloads.
  • Build automated workflows to detect, isolate, repair, and remediate unhealthy infrastructure while minimizing service disruption.
  • Manage fleet lifecycle administration including firmware updates, driver management, OS patching, and configuration drift identification.
  • Translate NVIDIA reference architectures into production operating practices, runbooks, and measurable operational standards.
  • Develop reusable tools, automation scripts, and implementation guides for consistent application across multiple partner environments.
  • Define operational health metrics, SLOs, and acceptance criteria to provide clear insight into infrastructure reliability.

What we're looking for

  • BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.
  • 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles.
  • Strong experience operating Linux-based distributed systems and cloud infrastructure in production environments.
  • Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments.
  • Strong understanding of production observability including metrics, logging, alerting, dashboards, health checks, and service level agreements.
  • Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
  • Strong networking fundamentals and experience troubleshooting complex distributed systems across compute, network, and storage layers.
  • Programming and automation experience using Python, Go, shell scripting, or similar languages.

More like this

Similar roles

NCX Senior Engineer

Nvidia

Remote (Santa Clara, CA) +1 8 days ago $184,000$287,500
Kubernetes Python Go PyTorch TensorFlow CUDA MLOps Terraform Ansible Prometheus Grafana InfiniBand RoCE Triton NeMo CI/CD Linux Distributed Systems NVIDIA Reference Architectures
8+ yrs exp Remote

Senior Solutions Architect, Generative AI

Nvidia

Remote (Santa Clara, CA) +1 43 days ago $184,000$287,500
GPU InfiniBand RoCE RDMA NCCL NVLink NVSwitch Kubernetes Slurm Python Linux High-Performance Computing Distributed Systems Shell Scripting DCGM Nsight Systems
6+ yrs exp Remote

Senior GNC Engineer

Anduril Industries

Seattle, WA 102 days ago $191,000$253,000
GNC State Estimation Sensor Fusion C++ Python Computer Vision VBN VIO VSLAM HIL CI/CD Git Robotics Aerospace Engineering Embedded Systems Terrain Contour Matching (TERCOM)
8+ yrs exp

Senior Product Dev Rel Engineer

Nvidia

Santa Clara, CA 23 days ago $224,000$356,500
Python C++ CUDA Kubernetes Docker PyTorch InfiniBand NVLink Triton Ray vLLM Kubeflow Linux GPU Acceleration
10+ yrs exp Hybrid

Senior GNC Engineer, Space

Anduril Industries

Washington, DC 165 days ago $191,000$253,000
GNC MATLAB Simulink Python C++ Go Linux Orbital Mechanics Control Theory Computer Vision Hardware-in-the-Loop monte-carlo simulation Spacecraft Dynamics Flight Software

NX CAD Support Engineer

Anduril Industries

Costa Mesa, CA 18 days ago $129,000$171,000
Siemens NX SolidWorks Teamcenter PLM Jira Service Desk CAD CAM CAE IT Helpdesk Technical Support
2+ yrs exp