Senior Software Engineer, Agentic AI and Observability

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$168,000–$270,250 / yr
Posted
3 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $209k
This role $219k
$153k most similar roles pay here $283k

This role pays more than 56% of similar roles. Most pay $172,762–$246,150 — the shaded band above. At the midpoint, this role pays about $219k versus about $209k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 896 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 876 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Software Engineer, Agentic AI and Observability

As a Senior Software Engineer, Agentic AI and Observability on the BizApps SRE team, you will advance the Agentic AI Factory model to accelerate the development, deployment, and operation of enterprise-wide AI applications. You will build and implement observability solutions for agentic AI systems, including metrics, tracing, logging, and alerting while developing deployment pipelines and reliability tools. Your daily work involves partnering with application teams to define SLOs and SLIs, monitoring LLM-based workflows for performance and cost, and identifying gaps like hallucination detection. You will utilize Python, Kubernetes, Docker, and observability platforms such as Datadog, OpenTelemetry, Grafana, or Prometheus. The role focuses on the technical challenge of ensuring reliability in multi-step reasoning chains and tool orchestration while managing complex distributed systems within a production environment to ensure high-quality AI service delivery.

What does a Software Engineer earn in California?

Median $214000 from 775 postings across 63 companies.

See salary data

What you'll do

  • Build and implement observability solutions including metrics, tracing, logging, and alerting for agentic AI applications.
  • Develop and maintain deployment pipelines and reliability tools for the Agentic AI Factory model.
  • Define SLOs and SLIs with application teams to ensure production readiness for new agent deployments.
  • Instrument LLM-based workflows to monitor performance, cost, token usage, and multi-step reasoning chains.
  • Lead incident response, root cause analysis, and reliability improvements for BizApps AI services.
  • Identify and resolve observability gaps such as hallucination detection and orchestration failure tracing.
  • Integrate reliability measures across the development lifecycle in collaboration with platform, data, and application teams.

What we're looking for

  • BS or MS in Computer Science, Software Engineering, or a related field (or equivalent experience).
  • 8+ years of experience and strong software engineering skills in Python.
  • Experience with modern CI/CD practices.
  • Hands-on experience with container orchestration (Kubernetes, Docker) and cloud infrastructure.
  • Experience with observability and monitoring platforms such as Datadog, OpenTelemetry, Grafana, or Prometheus.
  • SRE / DevOps approach to owning production systems, automating toil, and building for reliability.
  • Proven track record of debugging complex distributed systems and excellent problem-solving skills.
  • Experience with LLM orchestration frameworks, ML/AI production environments, infrastructure-as-code, or open-source contributions (preferred).

More like this

Similar roles

Senior Software Engineer, AIOps and Observability

Nvidia

Santa Clara, CA 42 days ago $200,000$322,000
AIOps Observability Prometheus Victoria Metrics Vector Loki Grafana Alert Manager Clickhouse OpenTelemetry BigPanda PagerDuty Datadog Kubernetes Nomad Docker Microservices NATS Kafka Go Python Java C# Machine Learning Generative AI LLMs
10+ yrs exp

Senior Site Reliability Engineer, AIOPs

Nvidia

Santa Clara, CA 122 days ago $148,000$235,750
Kubernetes Python Bash Terraform Helm CI/CD infrastructure-as-code Prometheus Grafana Kafka Pulsar Flink Spark ClickHouse Linux Distributed Systems Microservices SRE AIOps
5+ yrs exp

Senior Software Engineer, Agentic Systems

Adobe

San Jose, CA 64 days ago $183,300$265,350
LLM RAG Python PyTorch TensorFlow Kubernetes Docker AWS Transformers Diffusion Models GANs CLIP MLLMs NVIDIA Triton TorchServe ONNX CUDA Distributed Systems
5+ yrs exp

Senior Software Engineer, Agentic AI Tools

F5 Inc

Remote (San Jose, CA) 10 days ago $166,100$249,100
Python TypeScript Go LangChain LangGraph CrewAI LlamaIndex RAG Prompt Engineering MLflow Kubeflow BentoML BeautifulSoup Scrapy NGINX API Vector Databases Cybersecurity
8+ yrs exp Remote Hybrid