Senior Infrastructure Engineer, AI, Automation, Observability and Monitoring

Nvidia

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$200,000–$322,000 / yr
Posted
2 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $185k
This role $261k
$111k most similar roles pay here $345k

This role pays more than 96% of similar roles. Most pay $144,350–$226,350 — the shaded band above. At the midpoint, this role pays about $261k versus about $185k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Infrastructure Engineer, AI, Automation, Observability and Monitoring

As a Senior Infrastructure Engineer - AI, Automation, Observability and Monitoring, you will join a dynamic team to design, develop, and deploy world-class software solutions while mentoring junior engineers. You will architect systems that meet ambitious performance, scalability, and reliability requirements while advocating for standard methodologies like code reviews, testing, and continuous integration. Your daily work involves agentizing workflows and building agents for various domains of infrastructure using brand new technologies. The role requires extensive experience in C++, Python, and Java to solve complex engineering challenges. You will specifically utilize Claude and Codex to write agents, skills, and harnesses while coding integrations to tools like BigPanda, ITMP, and Prometheus. This position focuses on the technical challenge of automating workflows and implementing software solutions that address complex business problems through advanced software engineering techniques.

What you'll do

  • Design, develop, and deploy high-performance software solutions that meet scalability and reliability requirements.
  • Agentize workflows by building specialized agents for various domains of infrastructure.
  • Develop agents, skills, and harnesses using technologies like Claude and Codex.
  • Code integrations to monitoring and management tools such as BigPanda, ITMP, and Prometheus.
  • Implement automation and software delivery processes to solve complex business problems.
  • Advocate for standard development methodologies including code reviews, testing, and continuous integration.
  • Mentor junior engineers and provide technical leadership to foster a culture of growth.

What we're looking for

  • 12+ years of experience in software engineering roles with a focus on team leadership and project management.
  • Proven track record of crafting and implementing software solutions to solve complex business problems.
  • Expertise in programming languages including C++, Python, and Java.
  • Experience writing agents, skills, and harnesses using Claude and Codex.
  • Experience coding integrations to tools such as BigPanda, ITMP, and Prometheus.
  • Proven track record with automation and software development and delivery.
  • Experience working with Fortune 500 companies across diverse industries.
  • Master’s or Ph.D. in a relevant field (e.g., Computer Science, Software Engineering) or equivalent experience.
  • Proficiency in AI, machine learning, and data-driven solutions (preferred).

More like this

Similar roles

Senior Infrastructure Engineer

Wells Fargo

Concord, CA +2 8 days ago $104,000$168,000
Linux Apache Tomcat F5 LTM AVI AppDynamics Splunk Grafana Glassbox Dynatrace CI/CD Jenkins uDeploy Python PowerShell Ansible Salt Kafka Redis MongoDB Cassandra Kubernetes OpenShift Perl Shell JIRA Agile Kanban ITIL
4+ yrs exp Hybrid

AI Infrastructure Engineer

Blackrock

New York, NY 65 days ago $162,000$215,000
AWS Azure GCP Terraform Ansible CloudFormation Bicep Python Java Golang CI/CD MLOps Microservices Infrastructure as Code
5+ yrs exp Hybrid

AI Infrastructure Engineer

Blackrock

New York, NY 58 days ago $162,000$215,000
AWS Azure GCP Terraform Ansible CloudFormation Bicep Python Java Golang CI/CD MLOps Microservices Infrastructure as Code
5+ yrs exp Hybrid

AI Infrastructure Engineer

Fortinet

New York, NY 14 days ago $215,000$350,000
Linux GPU Docker Kubernetes Python Bash CI/CD KVM FortiGate FortiManager FortiAnalyzer Monitoring Logging Alerting Networking Virtualization Performance Testing Benchmarking Automation

AI Infrastructure Engineer

Amd

San Jose, CA 7 days ago $204,000$306,000
Kubernetes Terraform GitOps ArgoCD Flux Helm Prometheus Grafana Loki CSI CNI Slurm PyTorch vLLM SGLang Infrastructure as Code
Hybrid