Distinguished Software Engineer AI/ML Engineer Agentic Systems & Site Reliability Engineering

Walmart

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Sunnyvale, CA
Salary
$169,000–$338,000 / yr
Posted
78 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $202k
This role $254k
$139k most similar roles pay here $359k

This role pays more than 80% of similar roles. Most pay $162,000–$242,850 — the shaded band above. At the midpoint, this role pays about $254k versus about $202k for comparable roles.

Based on 240 similar postings.

Employer

About Walmart

Walmart Inc. is the world''s largest retailer by revenue, operating a chain of hypermarkets, discount department stores, and grocery stores, as well as a growing e-commerce presence through Walmart.com. Industry: General Merchandise & Grocery Retail

Walmart currently has 184 open roles on FindRole.

Listed pay typically runs $117,000–$234,000 across 180 roles with salary data.

Most-posted roles

View all roles at Walmart

At a glance

TL;DR · Distinguished Software Engineer AI/ML Engineer Agentic Systems & Site Reliability Engineering

(USA) Distinguished, Software Engineer-AI/ML Engineer - Agentic Systems & Site Reliability Engineering serves as a technical leader within the Site Reliability Engineering organization. This role focuses on architecting and implementing next-generation agentic AI systems and intelligent automation to ensure mission-critical reliability across e-commerce, supply chain, and in-store technologies. You will build multi-agent orchestration platforms, ML-driven anomaly detection systems, and self-healing infrastructure to automate incident response and capacity planning. The role requires expertise in machine learning frameworks like TensorFlow and PyTorch, agentic AI systems, LLMs, and cloud environments including Azure, GCP, and AWS. Key tools include Kubernetes, Docker, Terraform, Prometheus, Grafana, and the ELK stack. You will solve complex problems regarding system availability and performance by transforming traditional SRE practices into autonomous, predictive systems that proactively resolve issues across a vast, high-traffic retail technology ecosystem.

What you'll do

  • Architect and develop advanced agentic AI systems to automate complex reliability engineering workflows and predictive failure analysis.
  • Design multi-agent orchestration platforms for automated incident response, capacity planning, and performance optimization across all business units.
  • Build intelligent observability and monitoring systems using ML-driven anomaly detection and autonomous resolution capabilities.
  • Develop self-healing infrastructure platforms that use AI to predict and resolve system issues before they impact customers.
  • Create tools and automation to improve reliability, latency, availability, and scalability for mission-critical distributed systems.
  • Define and implement SLIs and SLOs with service owners to ensure all critical systems meet performance standards.
  • Build MLOps and AIOps platforms to enable continuous learning and automated optimization of reliability engineering systems.
  • Provide technical mentorship and guidance on AI/ML for reliability and platform engineering best practices to internal teams.

What we're looking for

  • Bachelor's degree in Engineering, Computer Science, or a related field with 6 years of experience in software engineering.
  • Alternatively, 8 years of experience in software engineering or a related area is accepted.
  • Master's degree in Computer Science or a related field with 4 years of experience in software engineering is preferred.
  • 12+ years of hands-on experience in Site Reliability Engineering, AI/ML Engineering, or Platform Engineering is required for the senior role.
  • Expert-level experience in machine learning algorithms, deep learning frameworks like TensorFlow and PyTorch, and production ML system deployment.
  • Advanced experience with agentic AI systems, including multi-agent frameworks, LLM-based agents, and autonomous decision-making platforms.
  • Comprehensive SRE expertise including service management, performance/capacity engineering for AI/ML, and large-scale distributed systems.
  • Expert-level cloud engineering skills in Azure, GCP, or AWS with knowledge of Kubernetes, Docker, and Infrastructure as Code.

More like this

Similar roles

Distinguished Software Engineer

Walmart

Bentonville, AR 85 days ago $130,000$260,000
AI Cloud-native CI/CD DevOps Observability Software Development Lifecycle Automation Gen AI Incident Management Developer Platforms
6+ yrs exp

Distinguished Software Engineer

Walmart

Sunnyvale, CA 95 days ago $169,000$338,000
Software Architecture Distributed Systems Cloud-native AI Agents Telemetry Root Cause Analysis Security
6+ yrs exp

Distinguished AI Engineer, Agentic AI Platform

Capital One Financial

Remote (San Francisco, CA) +4 46 days ago $269,100$307,200
Python Go Scala Java C++ C# LLM RAG LangGraph AutoGen Semantic Kernel CrewAI LlamaIndex Helm AWS Cloud Platform Azure SageMaker Vertex AI
8+ yrs exp Remote

Senior Agentic AI/ML Engineer

General Dynamics

Arlington, VA 73 days ago $199,750$270,250
Python Machine Learning LLM RAG Agentic AI CI/CD Docker Kubernetes Git PyTorch TensorFlow scikit-learn LangChain LangGraph PostgreSQL RESTful APIs MLOps LLMOps
10+ yrs exp

Senior Software Engineer, Agentic Systems

Adobe

San Jose, CA 64 days ago $183,300$265,350
LLM RAG Python PyTorch TensorFlow Kubernetes Docker AWS Transformers Diffusion Models GANs CLIP MLLMs NVIDIA Triton TorchServe ONNX CUDA Distributed Systems
5+ yrs exp