Software Platform Support Engineer, GPU Cloud

Nvidia

Confirmed live today High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CADurham, NC
Employment
Full-time
Posted
21 days ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $196k
$139k most similar roles pay here $252k

This listing doesn't post a salary. Most similar roles pay $161,300–$231,150.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1150 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 912 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Software Platform Support Engineer, GPU Cloud

The Software Platform Support Engineer - GPU Cloud joins the NVIDIA DGX Cloud organization to provide Tier 1 support for complex cloud platforms while partnering with internal customers. You will troubleshoot issues, investigate root causes, and create documentation in a fast-moving environment. Daily responsibilities include defining operational workflows like runbooks and escalation paths, filing bugs with Site Reliability teams, and building tools to improve visibility. The role requires expertise in Linux, Kubernetes, and major cloud providers including AWS, Azure, OCI, and GCP. You must possess skills in infrastructure, networking, storage technologies, and DevOps scripting. Candidates should have experience with distributed software systems and a background in troubleshooting complex environments. This position specifically addresses the technical challenges of supporting GPU workloads, MLOps workflows, and distributed training systems within a high-performance computing context.

What you'll do

  • Provide Tier 1 support for complex cloud platforms across compute, storage, and networking environments.
  • Triage and investigate the root cause of customer issues to determine necessary escalations.
  • Define and improve operational workflows including runbooks, escalation paths, and support processes.
  • File bugs and report technical issues while collaborating closely with the Site Reliability team.
  • Build internal tooling to improve visibility and streamline the customer support process.
  • Create documentation to help users troubleshoot common issues independently in a fast-moving environment.
  • Provide feedback to engineering teams based on user workloads to develop better solutions.
  • Participate in an on-call rotation to provide support for production systems.

What we're looking for

  • BS/MS degree in Computer Science or related fields (or equivalent experience).
  • 5+ years of experience supporting distributed software systems and end-user software platforms.
  • Experience with Linux operating systems.
  • Experience with Kubernetes, AWS, Azure, OCI, and GCP.
  • Background in infrastructure, networking, storage, and DevOps scripting/tooling.
  • Understanding of data storage technologies including databases, file, block, and blob.
  • Customer service/support experience and strong troubleshooting and communication skills.
  • Experience with MLOps workflows, ML infrastructure, GPU workloads, or SLURM/HPC (preferred).

More like this

Similar roles