Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

Confirmed live 2 days ago Low trust

Quick summary

Work type
On-site
Location
New York, NYMcLean, VACambridge, MASan Jose, CA
Salary
$197,300–$225,100 / yr
Posted
102 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $219k
This role $211k
$175k most similar roles pay here $272k

This role pays less than 53% of similar roles. Most pay $192,050–$246,150 — the shaded band above. At the midpoint, this role pays about $211k versus about $219k for comparable roles.

Based on 240 similar postings.

Employer

About Capital One Financial

Capital One Financial is a bank holding company specializing in credit cards, auto loans, banking, and savings products, known for its data-driven approach to consumer and commercial finance. Industry: Financial Services & Banking

Capital One Financial currently has 998 open roles on FindRole.

Listed pay typically runs $197,300–$225,100 across 992 roles with salary data.

Most-posted roles

View all roles at Capital One Financial

At a glance

TL;DR · Lead AI Engineer, FM Hosting, LLM Inference

Lead AI Engineer (FM Hosting, LLM Inference) joins the Intelligent Foundations and Experiences team to develop and deploy proprietary solutions that power core business functions. This role involves collaborating with cross-functional teams of engineers and scientists to design, test, and support critical software components including foundation model training, large language model inference, similarity search, guardrails, and observability. The engineer will implement state-of-the-art LLM optimization techniques to improve performance metrics like scalability, cost, latency, and throughput for production systems. Key technologies include PyTorch, Huggingface, VectorDBs, Nemo Guardrails, and AWS Ultraclusters. Candidates must possess proficiency in Python, Go, Scala, or Java, along with a strong foundation in engineering and mathematics. The role focuses on solving complex technical challenges related to high-performance AI infrastructure and the development of scalable, responsible machine learning systems for banking applications.

What you'll do

  • Design, develop, test, deploy, and support AI software components including foundation model training and LLM inference.
  • Implement similarity search, guardrails, model evaluation, and observability for production AI systems.
  • Develop state-of-the-art LLM optimization techniques to improve performance, scalability, cost, latency, and throughput.
  • Utilize a broad stack of technologies including AWS Ultraclusters, Huggingface, VectorDBs, and PyTorch.
  • Contribute to the technical vision and long-term roadmap of foundational AI systems.
  • Translate scientific research and novel techniques into practical production applications.
  • Solve complex, undefined problems by identifying root causes and providing clear technical solutions.

What we're looking for

  • Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms.
  • Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms.
  • At least 4 years of experience programming with Python, Go, Scala, or Java.
  • Experience deploying scalable and responsible AI solutions on cloud platforms such as AWS, Google Cloud, or Azure.
  • Experience designing, developing, delivering, and supporting AI services.
  • Experience developing AI and ML algorithms including LLM inference, similarity search, vector databases, and guardrails.
  • Experience applying state-of-the-art techniques to optimize training and inference software for hardware utilization, latency, throughput, and cost.
  • Ability to interpret scientific publications and apply novel research techniques in production environments.

More like this

Similar roles

Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

New York, NY +3 102 days ago $197,300$225,100
LLM Inference Python PyTorch Huggingface VectorDBs AWS Nemo Guardrails Go Scala Java C++ C# Similarity Search Machine Learning Model Evaluation Observability
4+ yrs exp

Senior Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

New York, NY +3 71 days ago $229,900$262,400
LLM Inference Python Go PyTorch Huggingface AWS VectorDBs Nemo Guardrails C++ Java Scala C# Machine Learning Similarity Search Model Evaluation Observability Optimization Techniques
6+ yrs exp

Senior Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

New York, NY +3 51 days ago $229,900$262,400
LLM Inference Python Go PyTorch Huggingface AWS VectorDBs Nemo Guardrails C++ Java Scala C# Machine Learning Similarity Search Model Evaluation Observability
6+ yrs exp

Senior Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

New York, NY +3 162 days ago $229,900$262,400
LLM Python Go PyTorch Huggingface AWS VectorDBs Nemo Guardrails C++ Java Scala C# Machine Learning LLM Inference Similarity Search Model Evaluation Observability
6+ yrs exp

Lead AI Engineer

Capital One Financial

New York, NY +4 21 days ago $197,300$225,100
LLM Python PyTorch Huggingface AWS VectorDBs Nemo Guardrails Go Scala Java C++ C# Machine Learning Similarity Search Model Inference SaaS
4+ yrs exp