Senior Software Engineer, AIOps and Observability

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$200,000–$322,000 / yr
Posted
42 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $182k
This role $261k
$96k most similar roles pay here $346k

This role pays more than 96% of similar roles. Most pay $150,000–$214,425 — the shaded band above. At the midpoint, this role pays about $261k versus about $182k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, AIOps and Observability

Senior Software Engineer, AIOps and Observability will join a team of engineers and product managers to design, develop, and deploy AIOps and observability platforms used to monitor and optimize millions of assets across cloud, on-prem, data center, supply chain, and edge environments. The role involves defining the technical roadmap, establishing standard methodologies, mentoring other engineers, and collaborating with data scientists to implement machine learning models for anomaly detection and root cause analysis. You will build agentic workflows and AI-native tools to process metrics, logs, traces, and events. Required skills include proficiency in Go, Python, Java, or C#, along with experience in Kubernetes, Docker, NATS, and Kafka. The candidate must be skilled with tools such as Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, and OpenTelemetry to manage large-scale distributed systems.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Design and develop AIOps and observability platforms including metrics, logs, traces, events, and dashboards.
  • Define the technical vision, roadmap, and standard methodologies for enterprise-wide observability initiatives.
  • Build agentic workflows and AI-native tools to automate anomaly detection, root cause analysis, and remediation.
  • Collaborate with data scientists to implement machine learning models for forecasting and automated debugging.
  • Develop scalable, distributed systems capable of handling high traffic and large volumes of telemetry data.
  • Establish observability standards and guidelines while researching new technologies to enhance user experience.
  • Provide peer reviews on performance, scalability, security, and correctness for other engineering team members.

What we're looking for

  • Bachelor's degree in computer science and engineering, a related field, or equivalent experience.
  • 12+ years of experience in product development and full stack engineering.
  • 5+ years of experience developing and operating observability platforms and solutions, preferably in cloud-native environments.
  • Proficiency in at least one programming language such as Go, Python, Java, or C#.
  • Experience with observability tools including Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, and OpenTelemetry.
  • Hands-on knowledge of AIOps tools like BigPanda, PagerDuty, and Datadog.
  • Experience with Kubernetes, Nomad, Docker, microservices architectures, and streaming services like NATS or Kafka.
  • Expertise in using machine learning, generative AI, or agentic AI frameworks to build predictive monitoring and automated remediation solutions.

More like this

Similar roles

Senior Software Engineer, Agentic AI and Observability

Nvidia

Santa Clara, CA 3 days ago $168,000$270,250
Python Kubernetes Docker CI/CD Datadog OpenTelemetry Grafana Prometheus Terraform Pulumi LangChain LlamaIndex Semantic Kernel GitOps SRE DevOps LLM Distributed Systems
8+ yrs exp

Software Development Engineer

Adobe

Seattle, WA +2 18 days ago $177,900$257,550
OpenTelemetry LLMs Python Go TypeScript JavaScript Kubernetes AWS CI/CD Jenkins Git Artifactory Splunk NewRelic DataDog Loki Prometheus Grafana AppDynamics
5+ yrs exp

Staff Software Engineer, Observability

Pinterest

Remote 7 days ago $177,185$364,795
Distributed Systems Data Engineering Observability OpenTelemetry Prometheus Grafana Kafka Flink Java Python Go Scala Kubernetes Service Mesh Time-series Databases Columnar Storage Stream Processing Machine Learning Anomaly Detection
7+ yrs exp Remote

Senior Site Reliability Engineer, AIOPs

Nvidia

Santa Clara, CA 122 days ago $148,000$235,750
Kubernetes Python Bash Terraform Helm CI/CD infrastructure-as-code Prometheus Grafana Kafka Pulsar Flink Spark ClickHouse Linux Distributed Systems Microservices SRE AIOps
5+ yrs exp

Senior Observability Automation Engineer

Q2

Austin, TX 11 days ago
C# Go Python Bash PowerShell Perl RESTful API CI/CD Infrastructure as Code Observability Grafana Splunk Cloud Infrastructure DevOps Linux Windows Event-driven Automation AI Platform
8+ yrs exp