Senior Software Engineer, Distributed Systems Engineer, EDA Infrastructure

Nvidia

Confirmed live 2 days ago High trust
Remote

Quick summary

Work type
Remote
Location
WAWestford, MAAustin, TXDurham, NC
Salary
$152,000–$241,500 / yr
Posted
3 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $163k
This role $197k
$102k most similar roles pay here $256k

This role pays more than 83% of similar roles. Most pay $149,350–$177,250 — the shaded band above. At the midpoint, this role pays about $197k versus about $163k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, Distributed Systems Engineer, EDA Infrastructure

Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure joins the team to build and scale infrastructure supporting Electronic Design Automation workloads. This role involves designing and building platforms to automate the provisioning, configuration, operation, and lifecycle management of large-scale GPU and CPU compute systems. The engineer will develop monitoring, health-management, and remediation systems while automating hardware deployment, firmware updates, and recovery workflows. Key responsibilities include integrating with workload schedulers and observability platforms using data from network, storage, and hardware diagnostics to ensure reliability for critical chip-design workloads. Candidates should possess strong programming skills in Go or Python, along with experience in distributed systems, Linux-based compute nodes, and infrastructure automation. The role addresses the technical challenge of managing complex, high-performance computing environments while reducing operational toil through production-quality software and automated system management across heterogeneous hardware.

What you'll do

  • Design and build platforms to automate the provisioning, configuration, and lifecycle management of large-scale GPU and CPU compute infrastructure.
  • Develop monitoring, health-management, and remediation systems to improve the reliability and availability of EDA compute environments.
  • Automate hardware deployment, operating-system configurations, firmware updates, and cluster enrollment workflows.
  • Build reliable services that integrate with workload schedulers, infrastructure management systems, and observability platforms.
  • Analyze hardware diagnostics, OS signals, and network telemetry to identify failures and restore unhealthy systems to service.
  • Perform incident response, root-cause analysis, and capacity planning for production services.
  • Develop scalable solutions for critical chip-design workloads by integrating with networking, storage, and hardware teams.

What we're looking for

  • 5+ years of software engineering or infrastructure engineering experience supporting large-scale production systems.
  • A BS in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent experience.
  • Strong programming experience in Go or Python including data structures, algorithms, testing, and software design.
  • Experience designing automation for distributed systems and large fleets of Linux-based compute nodes.
  • Understanding of performance, security, reliability, fault tolerance, state management, and data consistency in complex systems.
  • Experience with infrastructure automation, software deployment, observability, and operational recovery.
  • Strong communication skills and the ability to work effectively across teams and geographic regions.
  • A systematic approach to problem solving with a focus on reducing operational toil.

More like this

Similar roles

Senior Linux Systems Engineer, EDA Infrastructure

Nvidia

Remote (Westford, MA) +2 2 days ago $184,000$287,500
Linux Ansible Terraform Python Bash Go Docker Podman Kubernetes CI/CD GitLab CI GitHub Actions Jenkins AWS GCP LDAP NFS SDN Slurm LSF
8+ yrs exp Remote

Senior AI Infrastructure Engineer, EDA Infrastructure

Nvidia

Remote (Westford, MA) +2 8 days ago $184,000$287,500
Python Go TypeScript Java Telemetry Pipelines Metrics Logs Traces Observability SaaS ML Models AI Agent Frameworks CMDB incident-management Configuration Management
8+ yrs exp Remote

Senior Software Engineer, System Validation, EDA Infrastructure

Nvidia

Remote (Santa Clara, CA) +3 5 days ago $224,000$356,500
Cloud Infrastructure Distributed Systems DevOps SRE Platform Engineering Bare Metal as a Service ML/AI Infrastructure High-Performance Computing Workload Orchestration Hardware Multi-cloud Infrastructure Workflows
10+ yrs exp Remote

Senior Systems Software Engineer, EDA Infrastructure

Nvidia

Remote 3 days ago $184,000$287,500
Python Go Linux Containers Kubernetes Docker OpenStack Slurm Distributed Systems Infrastructure Automation Networking Storage Technologies Bare Metal as a Service NVIDIA GPU High-Performance Computing
8+ yrs exp Remote

Principal Software Engineer, Distributed Systems Engineer

Nvidia

Remote (Durham, NC) 78 days ago $272,000$431,250
Kubernetes GPU Go Python Slurm Bright Cluster Manager Distributed Systems Cluster Management Monitoring Data Structures Algorithms Systems Programming Network Telemetry Incident Management
10+ yrs exp Remote