Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

Confirmed live 2 days ago Low trust

Quick summary

Work type
On-site
Location
New York, NYMcLean, VACambridge, MASan Jose, CA
Salary
$197,300–$225,100 / yr
Posted
102 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $219k
This role $211k
$175k most similar roles pay here $272k

This role pays less than 53% of similar roles. Most pay $192,050–$246,150 — the shaded band above. At the midpoint, this role pays about $211k versus about $219k for comparable roles.

Based on 240 similar postings.

Employer

About Capital One Financial

Capital One Financial is a bank holding company specializing in credit cards, auto loans, banking, and savings products, known for its data-driven approach to consumer and commercial finance. Industry: Financial Services & Banking

Capital One Financial currently has 998 open roles on FindRole.

Listed pay typically runs $197,300–$225,100 across 992 roles with salary data.

Most-posted roles

View all roles at Capital One Financial

At a glance

TL;DR · Lead AI Engineer, FM Hosting, LLM Inference

Lead AI Engineer (FM Hosting, LLM Inference) joins the Intelligent Foundations and Experiences team to develop and deploy proprietary solutions that power core business functions. This role involves collaborating with cross-functional teams of engineers, research scientists, and product managers to build AI-powered products. The engineer will design, test, and support software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, and observability. Key responsibilities include implementing state-of-the-art LLM optimization techniques to improve performance metrics like scalability, cost, latency, and throughput for production systems. The role requires proficiency in Python, Go, Scala, or Java, along with experience using tools such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, and PyTorch. This position focuses on the technical challenge of building high-performance AI infrastructure to solve complex problems within the banking domain.

What you'll do

  • Design, develop, test, deploy, and support AI software components including foundation model training and LLM inference.
  • Implement similarity search, guardrails, model evaluation, experimentation, governance, and observability for production systems.
  • Develop state-of-the-art LLM optimization techniques to improve performance, scalability, cost, latency, and throughput.
  • Utilize a broad stack of technologies including AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, and PyTorch.
  • Contribute to the technical vision and long-term roadmap of foundational AI systems at Capital One.
  • Translate complex scientific publications into practical, high-performance production features for banking services.

What we're looking for

  • Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus 4 years of experience developing AI/ML algorithms.
  • Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus 2 years of experience developing AI/ML algorithms.
  • At least 4 years of experience programming with Python, Go, Scala, or Java.
  • Experience deploying scalable and responsible AI solutions on cloud platforms like AWS, Google Cloud, or Azure.
  • Experience designing, developing, delivering, and supporting AI services.
  • Experience developing AI and ML technologies including LLM inference, similarity search, vector databases, and guardrails.
  • Experience applying state-of-the-art techniques to optimize training and inference software for hardware utilization, latency, throughput, and cost.
  • Ability to understand scientific publications and apply novel research techniques in production environments.

More like this

Similar roles

Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

New York, NY +3 102 days ago $197,300$225,100
LLM Inference Python PyTorch Huggingface VectorDBs AWS Nemo Guardrails Go Scala Java C++ C# Similarity Search Machine Learning Model Evaluation Observability
4+ yrs exp

Senior Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

New York, NY +3 51 days ago $229,900$262,400
LLM Inference Python Go PyTorch Huggingface AWS VectorDBs Nemo Guardrails C++ Java Scala C# Machine Learning Similarity Search Model Evaluation Observability
6+ yrs exp

Senior Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

New York, NY +3 71 days ago $229,900$262,400
LLM Inference Python Go PyTorch Huggingface AWS VectorDBs Nemo Guardrails C++ Java Scala C# Machine Learning Similarity Search Model Evaluation Observability Optimization Techniques
6+ yrs exp

Senior Lead AI Engineer, FM Hosting, LLM Inference

Capital One Financial

New York, NY +3 162 days ago $229,900$262,400
LLM Python Go PyTorch Huggingface AWS VectorDBs Nemo Guardrails C++ Java Scala C# Machine Learning LLM Inference Similarity Search Model Evaluation Observability
6+ yrs exp

Lead AI Engineer

Capital One Financial

New York, NY +4 21 days ago $197,300$225,100
LLM Python PyTorch Huggingface AWS VectorDBs Nemo Guardrails Go Scala Java C++ C# Machine Learning Similarity Search Model Inference SaaS
4+ yrs exp