Director AI Engineering

Capital One Financial

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
New York, NYSan Francisco, CAMcLean, VACambridge, MASan Jose, CA
Salary
$244,700–$279,200 / yr
Employment
Full-time
Posted
5 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $240k
This role $262k
$172k $313k
below market most similar roles pay here above market

This role pays more than 62% of similar roles. Most pay $195,000–$285,086 — the blue band above. At the midpoint, this role pays about $262k versus about $240k for comparable roles.

Based on 240 similar postings.

Employer

About Capital One Financial

Capital One Financial is a bank holding company specializing in credit cards, auto loans, banking, and savings products, known for its data-driven approach to consumer and commercial finance. Industry: Financial Services & Banking

Capital One Financial currently has 1891 open roles on FindRole.

Listed pay typically runs $197,300–$225,100 across 1020 roles with salary data.

Most-posted roles

View all roles at Capital One Financial

At a glance

TL;DR · Director AI Engineering

The Director AI Engineering leads the Generative AI Training team to build a platform for scientists and engineers to train, fine-tune, and experiment with foundation models at scale. You will oversee the design, development, and operation of core systems including distributed training, reinforcement learning workflows, fair-share GPU scheduling, and self-service environments for model evaluation. Key responsibilities involve managing GPU capacity planning, cost governance, and establishing enterprise standards for Responsible AI. You will utilize a broad stack of technologies including AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, KServe, vLLM, Kubernetes, Kubeflow, and CUDA. The role solves the technical challenge of creating scalable, high-performance AI infrastructure that ensures efficient hardware utilization and reliable model deployment across various product areas while maintaining strict ethical and regulatory standards.

What you'll do

  • Oversee the design, development, and operation of distributed training and fine-tuning infrastructure for foundation models.
  • Manage fair-share GPU scheduling, job resilience, and utilization controls to ensure efficient hardware usage.
  • Provide self-service environments for teams to deploy, serve, and evaluate models using tools like KServe and vLLM.
  • Make build-vs-buy decisions for open-source and SaaS AI technologies including AWS Ultraclusters, Huggingface, and VectorDBs.
  • Own GPU capacity planning and cost governance by right-sizing clusters, instance types, and quotas.
  • Translate enterprise AI strategy into execution plans while establishing standards for Responsible AI and fairness metrics.
  • Scale AI engineering practices through shared infrastructure, reusable components, and unified observability frameworks.
  • Lead and mentor a team of engineers to attract top talent and foster a culture of continuous learning.

What we're looking for

  • Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus 8 years of experience developing AI and ML algorithms or technologies.
  • Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus 6 years of experience developing AI and ML algorithms or technologies.
  • At least 3 years of people leadership experience.
  • 5+ years of experience managing and leading an engineering team (preferred).
  • 7+ years of experience building and operating large-scale ML or GPU training infrastructure on cloud platforms (preferred).
  • Hands-on experience with distributed training at scale, including multi-node/multi-GPU jobs and proficiency in Python, Go, C++, or CUDA (preferred).
  • Experience with ML orchestration and scheduling stacks such as Kubernetes, Kubeflow, Kueue, Slurm, Ray, KServe, or vLLM (preferred).
  • Experience operating large GPU fleets with a focus on reliability, fault tolerance, utilization, and cost efficiency (preferred).

More like this

Similar roles

Director, AI Engineer

Capital One Financial

Remote (San Francisco, CA) +5 16 days ago $244,700–$279,200
Python C++ Java CUDA Golang PyTorch Huggingface AWS Google Cloud Azure VectorDBs LLM Inference Similarity Search Agentic AI Observability Governance Responsible AI
8+ yrs exp Remote

Director, AI Engineering

Capital One Financial

Remote (McLean, VA) +2 30 days ago
AI Machine Learning LLM PyTorch AWS Huggingface VectorDB Nemo Guardrails Observability Similarity Search Model Evaluation SaaS Open Source
8+ yrs exp Remote

Senior Director, AI Engineering, Agentic AI Platform

Capital One Financial

Remote (San Francisco, CA) +4 81 days ago
LLM PyTorch AWS VectorDB Huggingface Nemo Guardrails Similarity Search Model Evaluation Observability Machine Learning Google Cloud Azure
10+ yrs exp Remote

Senior Director, AI Engineering

Capital One Financial

San Francisco, CA +5 45 days ago
Machine Learning Python PyTorch AWS Huggingface VectorDBs Nemo Guardrails C++ C# Java Golang Similarity Search LLM Inference Observability
10+ yrs exp

Staff AI Engineer

Capital One Financial

Remote (San Francisco, CA) +3 41 days ago
Python Go Scala Java C++ C# PyTorch AWS Huggingface VectorDBs Nemo Guardrails LLM Inference Similarity Search Machine Learning Google Cloud Azure
8+ yrs exp Remote

Staff AI Engineer

Capital One Financial

Remote (San Francisco, CA) +3 19 days ago $244,700–$279,200
Python Go Scala CUDA Java C++ C# PyTorch AWS Huggingface VectorDBs LLM Inference Similarity Search Agentic Workflows Observability Model Evaluation Data Pipeline Governance Google Cloud Azure
8+ yrs exp Remote