Principal Software Engineer, E2E Performance and Goodput

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CAAustin, TX
Salary
$272,000–$431,250 / yr
Posted
5 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $200k
This role $352k
$102k most similar roles pay here $467k

This role pays more than 99% of similar roles. Most pay $174,600–$226,350 — the shaded band above. At the midpoint, this role pays about $352k versus about $200k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Principal Software Engineer, E2E Performance and Goodput

Principal Software Engineer, E2E Performance and Goodput — CSP Engagements joins the CSP Engagements team as a technical focal point for end-to-end performance. This role involves collaborating with engineering teams from key CSP and hyperscale customers to ensure they achieve specific performance targets on NVIDIA platforms. You will lead work streams to characterize platform performance, gather workload-specific feedback to influence optimization priorities across CUDA, NCCL, driver, and firmware teams, and validate performance in customer-representative configurations. Key responsibilities include updating open-source tools like STREAM, GPU Burn, and GPU BLAST while analyzing cross-CSP patterns to identify systemic improvements. Required skills include expertise in GPU workload profiling using Nsight Systems, Nsight Compute, and DCGM metrics; proficiency in Python and pandas for data analysis; and deep knowledge of distributed training dynamics, including collective efficiency and memory bandwidth utilization.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Act as the technical focal point for end-to-end performance with key CSP and hyperscale customers.
  • Lead performance characterization work streams to educate customers on platform expectations, profiling methods, and tuning options.
  • Synthesize customer feedback to identify performance gaps and advocate for optimization priorities within internal engineering teams.
  • Update and validate open-source performance and stress tools for the latest NVIDIA rack-scale systems and architectures.
  • Analyze cross-CSP performance patterns to identify software or configuration issues causing performance gaps.
  • Ensure customer performance and validation tooling reflects current GPU capabilities and memory hierarchy changes.
  • Coordinate with customers to ensure profiling infrastructure and benchmark harnesses are ready before deployment milestones.
  • Define test strategies and tooling requirements for both internal certification and customer acceptance.

What we're looking for

  • 15+ years of experience in systems performance engineering, ideally in GPU/HPC/ML infrastructure.
  • BS or MS in Computer Science, Computer Engineering, or a related field (or equivalent experience).
  • Proficiency in GPU workload profiling using tools like Nsight Systems, Nsight Compute, and DCGM metrics.
  • Understanding of distributed training performance dynamics including computation/communication overlap and memory bandwidth utilization.
  • Knowledge of how the full software stack impacts performance, including driver overhead, firmware power management, and scheduling.
  • Strong data analysis and visualization skills using Python, pandas, and dashboards.
  • Proficiency in statistical methods for performance analysis such as regression detection and A/B comparison at scale.
  • Ability to communicate complex technical findings to both deep technical audiences and executive leadership.

More like this

Similar roles

Principal Software Engineer, Core Infrastructure

Oracle

Seattle, WA 107 days ago $114,600$234,600
Distributed Systems Distributed Databases Java Oracle Cloud Infrastructure (OCI) Transaction Processing Data-plane Architecture Caching Capacity Planning Observability Automation
6+ yrs exp

Principal Core Infrastructure Engineer

Oracle

Nashville, TN 35 days ago $114,600$234,600
Oracle Cloud Infrastructure Microservices Firmware SmartNIC ILOM NVIDIA AMD Intel Automation Pipelines Observability Hardware-Software Integration
8+ yrs exp

Principal Software Engineer, Core Infrastructure

Oracle

Nashville, TN 27 days ago $114,600$234,600
Distributed Systems System Design Infrastructure as Code (IaC) Data Processing Telemetry Monitoring Alerting Fault Injection rate‑limiting Encryption Cloud Infrastructure Scalability Automation

Principal Software Engineer, Core Infrastructure

Oracle

Santa Clara, CA +1 27 days ago $114,600$234,600
Distributed Systems Infrastructure as Code (IaC) Data Processing Telemetry Monitoring Alerting rate‑limiting Fault Injection Encryption Automation Scalability
3+ yrs exp

Principal Software Engineer, Core Infrastructure

Oracle

Nashville, TN 27 days ago $114,600$234,600
Distributed Systems Infrastructure as Code (IaC) Data Processing Telemetry Monitoring Alerting Fault Injection rate‑limiting Encryption Automation Scalability Data Replication Multi-tenant Environments
3+ yrs exp

Principal Software Engineer, Core Infrastructure

Oracle

Nashville, TN 27 days ago $114,600$234,600
Distributed Systems Infrastructure as Code (IaC) Cloud Infrastructure Data Processing System Design Telemetry Monitoring Alerting Fault Injection rate‑limiting Encryption Automation Scalability
6+ yrs exp