Senior AI Infrastructure Engineer, EDA Infrastructure

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Westford, MAAustin, TXDurham, NC
Salary
$184,000–$287,500 / yr
Posted
8 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $190k
This role $236k
$125k most similar roles pay here $305k

This role pays more than 77% of similar roles. Most pay $144,350–$235,750 — the shaded band above. At the midpoint, this role pays about $236k versus about $190k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior AI Infrastructure Engineer, EDA Infrastructure

As a Senior AI Infrastructure Engineer - EDA Infrastructure, you will join the team building systems, tooling, and data infrastructure to support GPU cloud services. You will develop scalable telemetry pipelines for metrics, logs, traces, and events across on-premise, CSP, and NCP clusters while establishing common instrumentation and storage patterns. Your daily work involves standardizing incident management workflows, creating automated reporting tools to reduce manual toil, and maintaining physical hardware and software catalogs as trusted sources of truth. You will build integration pipelines for service ownership and provide self-service discovery capabilities. The role requires proficiency in Python, Go, TypeScript, or Java, along with expertise in infrastructure platform engineering. You will solve complex problems regarding service visibility, automated troubleshooting, and inventory management within the specialized domain of EDA infrastructure and large-scale AI platform operations.

What you'll do

  • Build and operate scalable telemetry pipelines for metrics, logs, traces, and events across multiple cluster types.
  • Establish common instrumentation, collection, storage, and access patterns to ensure consistent telemetry consumption.
  • Deliver dashboards, alerting systems, and analysis tools to improve service visibility and troubleshooting.
  • Automate incident management, maintenance, and on-call workflows to reduce manual toil and improve responsiveness.
  • Integrate operational data and lifecycle signals to enhance ownership, escalation, and post-incident learning.
  • Maintain physical hardware and software catalogs as trusted sources of truth for infrastructure inventory.
  • Create consistent data models and integration pipelines connecting clusters, hardware, services, and teams.
  • Provide self-service discovery capabilities for engineers to identify service ownership and support procedures.

What we're looking for

  • Bachelor's degree in Computer Science, Computer Engineering, or a related technical field, or equivalent experience.
  • 8+ years of experience in infrastructure security, platform engineering, or security tooling.
  • Proficiency in one or more programming languages such as Python, Go, TypeScript, or Java.
  • Strong understanding of software and infrastructure principles applied in production environments.
  • Ability to lead cross-functional initiatives across engineering, product, finance, and security teams.
  • Experience building and operating incident-management processes with internal or external SaaS tools.
  • Experience deploying and maintaining ML models in production systems with familiarity with AI agent frameworks.
  • Experience building modern observability platforms for metrics, logs, traces, and profiling.

More like this

Similar roles

AI Infrastructure Engineer

Blackrock

New York, NY 64 days ago $162,000$215,000
AWS Azure GCP Terraform Ansible CloudFormation Bicep Python Java Golang CI/CD MLOps Microservices Infrastructure as Code
5+ yrs exp Hybrid

AI Infrastructure Engineer

Blackrock

New York, NY 57 days ago $162,000$215,000
AWS Azure GCP Terraform Ansible CloudFormation Bicep Python Java Golang CI/CD MLOps Microservices Infrastructure as Code
5+ yrs exp Hybrid

AI Infrastructure Engineer

Fortinet

New York, NY 13 days ago $215,000$350,000
Linux GPU Docker Kubernetes Python Bash CI/CD KVM FortiGate FortiManager FortiAnalyzer Monitoring Logging Alerting Networking Virtualization Performance Testing Benchmarking Automation