Senior Systems Software Engineer, Observability and Telemetry Platform

Nvidia

Confirmed live today High trust
Remote

Quick summary

Work type
Remote
Location
CA
Salary
$152,000–$241,500 / yr
Posted
3 days ago
Freshness
Confirmed live today

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $192k
This role $197k
$141k most similar roles pay here $252k

This role pays more than 50% of similar roles. Most pay $155,900–$228,950 — the shaded band above. At the midpoint, this role pays about $197k versus about $192k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 929 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 915 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Systems Software Engineer, Observability and Telemetry Platform

As a Senior Systems Software Engineer, Observability and Telemetry Platform, you will join the engineering team to compose, build, and maintain large-scale production systems focused on high efficiency and availability. You will design, implement, and support operational aspects of an observability and telemetry collection platform, focusing on real-time monitoring, logging, alerting, and performance at scale. Your daily work involves managing the full service lifecycle from inception through deployment, performing capacity management, and automating routine tasks to improve reliability and velocity. The role requires expertise in distributed systems design, Linux, networking, and containers. You will utilize tools such as Kubernetes, OpenStack, Docker, Grafana, Prometheus, and OpenTelemetry while coding in Python, Go, Perl, or Ruby. This position addresses the critical challenge of ensuring maximum uptime for GPU cloud services while optimizing system health and performance.

What you'll do

  • Design and implement large-scale observability and telemetry collection platforms for real-time monitoring, logging, and alerting.
  • Manage the full lifecycle of services from initial design and consulting through deployment and refinement.
  • Monitor and maintain system health by measuring availability, latency, and performance metrics.
  • Scale systems sustainably through automation and proactive performance tuning to eliminate manual work.
  • Conduct blameless postmortems and practice sustainable incident response for production systems.
  • Participate in an on-call rotation to provide support for live production environments.
  • Develop software tools, platforms, and frameworks to support system capacity management and launch reviews.

What we're looking for

  • BS degree in Computer Science or a related technical field involving coding, or equivalent experience.
  • 5+ years of experience with Infrastructure automation and distributed systems design.
  • 5+ years of experience delivering foundational infrastructure and observability platforms.
  • Experience designing and developing tools for running large scale private or public cloud systems in production.
  • In-depth knowledge of Linux, Networking, and Containers.
  • Experience in one or more of the following: Python, Go, Perl, or Ruby.
  • Experience using or running large private and public cloud systems based on Kubernetes, OpenStack, and Docker (preferred).
  • Experience running Grafana, OpenTelemetry, Prometheus, and similar observability focused tools (preferred).

More like this

Similar roles

Senior Software Engineer, Observability

Apple Inc

Cary, NC 129 days ago
OpenTelemetry Grafana Datadog Kotlin Go Python Java Kubernetes Prometheus CI/CD Terraform Pulumi LLMs NoSQL Distributed Systems SRE API Design Infrastructure-as-Code
7+ yrs exp

Staff Software Engineer, Observability

Pinterest

Remote 10 days ago $177,185$364,795
Distributed Systems Data Engineering Observability OpenTelemetry Prometheus Grafana Kafka Flink Java Python Go Scala Kubernetes Service Mesh Time-series Databases Columnar Storage Stream Processing Machine Learning Anomaly Detection
7+ yrs exp Remote

Senior Software Engineer, AIOps and Observability

Nvidia

Santa Clara, CA 45 days ago $200,000$322,000
AIOps Observability Prometheus Victoria Metrics Vector Loki Grafana Alert Manager Clickhouse OpenTelemetry BigPanda PagerDuty Datadog Kubernetes Nomad Docker Microservices NATS Kafka Go Python Java C# Machine Learning Generative AI LLMs
10+ yrs exp

Senior Systems Software Engineer, EDA Infrastructure

Nvidia

Remote 2 days ago $184,000$287,500
Python Go Linux Containers Kubernetes Docker OpenStack Slurm Distributed Systems Infrastructure Automation Networking Storage Technologies Bare Metal as a Service NVIDIA GPU High-Performance Computing
8+ yrs exp Remote

Senior Engineer, Observability Platform

CVS Health

Remote (RI) +4 28 days ago $83,430$222,480
OpenTelemetry Go Python Java Spring Boot Kubernetes Docker Terraform CloudFormation Helm Kustomize Grafana Loki Tempo Mimir PostgreSQL MySQL Kafka Istio Envoy CI/CD
5+ yrs exp Remote