Executive Director Machine Learning Engineer MLOps

JPMorgan Chase

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Palo Alto, CA
Posted
38 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

How this pay compares to similar roles

Similar $229k
$174k most similar roles pay here $289k

This listing doesn't post a salary. Most similar roles pay $202,800–$255,062.

Based on 240 similar postings.

Employer

About JPMorgan Chase

JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.

JPMorgan Chase currently has 1117 open roles on FindRole.

Listed pay typically runs $186,160–$215,000 across 7 roles with salary data.

Most-posted roles

View all roles at JPMorgan Chase

At a glance

TL;DR · Executive Director Machine Learning Engineer MLOps

As an Executive Director Machine Learning Engineer-MLOps on the Recommendation Engine team, you will collaborate with Data Scientists to build and deploy machine learning models within a high-throughput, low-latency environment. You will be responsible for implementing fine-tuning and reinforcement learning algorithms on large compute clusters, managing real-time and batch model serving systems, and performing hyper-parameter tuning at scale. Your daily work involves building distributed training pipelines on GPU-enabled clusters, optimizing vector databases, and establishing monitoring and observability pipelines. You will utilize Python, AWS, vLLM, Ray, and CUDA to deploy open-weight large language models while applying quantization techniques like PTQ and AWQ. The role focuses on the technical challenges of serving transformer-based models and developing infrastructure for personalization and insights within a complex banking ecosystem involving travel, merchant offers, and dining services.

What you'll do

  • Build and maintain robust pipelines for distributed training on GPU-enabled clusters.
  • Develop and manage high-volume real-time and batch inference systems.
  • Implement quantization techniques to deploy large language models on modern serving stacks.
  • Manage and optimize vector databases to support advanced AI applications.
  • Establish comprehensive monitoring and observability pipelines for system health and performance.
  • Perform hyper-parameter tuning at scale for model development and deployment.
  • Integrate new technologies into existing infrastructure to improve production systems.

What we're looking for

  • BS in Computer Science or related field with 10+ years experience, or MS degree with 6+ years experience.
  • Extensive experience in Python and cloud computing, specifically within the AWS environment.
  • Understanding of quantization techniques like PTQ and AWQ to accelerate LLM inference on GPU architectures.
  • Solid understanding of Transformer models and the challenges involved in serving large transformer-based models.
  • Solid understanding of ML training, especially reinforcement learning algorithms such as GRPO and DAPO.
  • Experience in systems engineering fundamentals including caching, CUDA, autoscaling, high throughput, low latency, and x-region resilient applications.
  • Experience with monitoring and observability tools to monitor model input/output and feature statistics.
  • Experience with recommendation systems (preferred); experience with Ray, vLLM, RL libraries, or Kubernetes (preferred).

More like this

Similar roles

Lead Machine Learning Engineer, MLOps

JPMorgan Chase

Palo Alto, CA 72 days ago
MLOps Python AWS LLMs Quantization Ray vllm SGLang DuckDB Spark CUDA Docker Kubernetes ECS Airflow Kubeflow Vector Databases Distributed Training
6+ yrs exp

Senior Engineer, Machine Learning

Qualcomm

San Diego, CA 72 days ago $140,800$211,200
Python C++ Rust Go PyTorch TensorFlow LLM Transformer vLLM ONNX Kubernetes Docker CI/CD RAG Vector Databases OpenSearch Qdrant Quantization Distributed Computing Microservices
2+ yrs exp

Senior Engineer, Machine Learning

Qualcomm

San Diego, CA 72 days ago $140,800$211,200
Python C++ Rust Go PyTorch TensorFlow LLM Transformer vLLM ONNX Kubernetes Docker CI/CD RAG Vector Databases OpenSearch Qdrant Quantization Distributed Computing Microservices
2+ yrs exp

Director, Machine Learning Engineering

GEICO

Palo Alto, CA +4 98 days ago $150,000$300,000
AI ML RAG LLMs Generative AI Distributed Systems Platform Engineering Context Orchestration Memory Systems Retrieval Systems Real-time Inference Observability System Optimization Agent-based Systems
10+ yrs exp