Staff ML Software Engineer (L6)

Netflix

Confirmed live today High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Salary
$600,000–$1,066,000 / yr
Employment
Full-time
Posted
16 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $229k
This role $833k
$79k most similar roles pay here $1172k

This role pays more than 99% of similar roles. Most pay $203,850–$254,750 — the shaded band above. At the midpoint, this role pays about $833k versus about $229k for comparable roles.

Based on 240 similar postings.

Employer

About Netflix

Netflix is the world''s leading streaming entertainment service, offering a vast library of TV series, films, documentaries, and original content to subscribers in over 190 countries. Industry: Streaming Entertainment & Media

Netflix currently has 71 open roles on FindRole.

Listed pay typically runs $440,000–$725,000 across 59 roles with salary data.

Most-posted roles

View all roles at Netflix

At a glance

TL;DR · Staff ML Software Engineer (L6)

The Staff ML Software Engineer (L6) joins the Platform Systems team within AIMS Engineering to modernize the infrastructure supporting AI systems for recommendations, search, and personalized experiences. This role focuses on owning observability, cost, and platform subsystems to support next-generation AI workflows. You will design and operate subsystems for anomaly detection, root cause analysis, and operational automation while building observability primitives for training pipelines, serving latency, and data quality. Key responsibilities include driving cost optimization across training and serving infrastructure and architecting reliability improvements to reduce toil. The role requires deep Python expertise, proficiency in Scala or Java, and experience with distributed systems, batch processing, and real-time serving. You will build orchestration logic for agentic architectures, including memory, trace, and replay pipelines for complex model systems.

What you'll do

  • Design and operate subsystems for observability, evaluation, and tooling for next-generation ML architectures.
  • Build observability primitives to provide visibility into model behavior, training pipeline health, and serving latency.
  • Identify and drive cost optimization across AI training and serving infrastructure through automated frameworks.
  • Architect reliability improvements to reduce toil and improve on-call ergonomics across the AI/ML stack.
  • Contribute to the target architecture and migration path for the modernized AIMS AI/ML stack.
  • Evaluate emerging infrastructure patterns and model paradigms to develop a forward-looking platform roadmap.
  • Develop subsystems for advanced agentic architectures, including memory, trace, eval, and replay pipelines.

What we're looking for

  • Significant experience designing, building, and operating production AI/ML systems at scale, including training pipelines and model serving.
  • Hands-on experience building subsystems for advanced agentic architectures, such as memory, trace, eval, and replay pipelines.
  • Strong software engineering fundamentals with deep Python expertise and working proficiency in Scala or Java.
  • Proven track record of improving AI/ML system reliability, reducing infrastructure costs, and improving operational scalability.
  • Experience building observability and monitoring systems for AI/ML workloads across training, serving, and data pipelines.
  • Strong distributed systems background, including batch processing at scale and real-time serving infrastructure.
  • Ability to collaborate with partner teams to drive technical programs and build consensus without formal authority.
  • Familiarity with LLM evaluation, trace, or replay tooling (preferred).

More like this

Similar roles

Staff Machine Learning Engineer

Intuit

Mountain View, CA 44 days ago $230,000–$250,000
AI ML LLM RAG Python PyTorch TensorFlow Pandas NumPy LangChain AWS SageMaker Kubernetes Microservices JavaScript Java NoSQL System Design
8+ yrs exp

Staff Machine Learning Engineer, ML Platform

Braze

Austin, TX +1 22 days ago $184,000–$314,000
Machine Learning Python Ruby on Rails Kubernetes MongoDB Redis Kafka RabbitMQ Ray MLflow CI/CD Infrastructure as Code Distributed Systems Cloud Infrastructure Feature Stores Model Serving
8+ yrs exp Hybrid

Staff Machine Learning Engineer, ML Platform

Braze

New York City, NY 22 days ago $184,000–$314,000
Machine Learning Python Ruby on Rails Kubernetes CI/CD Infrastructure as Code MLflow Kafka RabbitMQ Celery Ray MongoDB Redis Distributed Systems Cloud Infrastructure Feature Stores Model Serving
8+ yrs exp Hybrid

Staff Machine Learning Engineer, ML Platform

Braze

Chicago, IL +1 22 days ago $184,000–$314,000
Machine Learning Python Ruby on Rails Kubernetes MongoDB Redis Kafka RabbitMQ Ray MLflow CI/CD Infrastructure as Code Distributed Systems Cloud Infrastructure Feature Stores Model Serving
8+ yrs exp Hybrid

Staff Machine Learning Engineer, ML Platform

Braze

San Francisco, CA +1 22 days ago $184,000–$314,000
Machine Learning Python Ruby on Rails Kubernetes MongoDB Redis Kafka RabbitMQ Ray MLflow CI/CD Infrastructure as Code Distributed Systems Cloud Infrastructure Feature Stores Model Serving
8+ yrs exp Hybrid