AI Engineer 5, FM Hosting, LLM Inference

Capital One Financial

Confirmed live today High trust

Quick summary

Work type
On-site
Location
New York, NYSan Francisco, CAMcLean, VACambridge, MASan Jose, CA
Salary
$229,900–$262,400 / yr
Employment
Full-time
Posted
2 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $231k
This role $246k
$203k $269k
below market most similar roles pay here above market

This role pays more than 75% of similar roles. Most pay $211,200–$250,100 — the blue band above. At the midpoint, this role pays about $246k versus about $231k for comparable roles.

Based on 240 similar postings.

Employer

About Capital One Financial

Capital One Financial is a bank holding company specializing in credit cards, auto loans, banking, and savings products, known for its data-driven approach to consumer and commercial finance. Industry: Financial Services & Banking

Capital One Financial currently has 1917 open roles on FindRole.

Listed pay typically runs $197,300–$225,100 across 1045 roles with salary data.

Most-posted roles

View all roles at Capital One Financial

At a glance

TL;DR · AI Engineer 5, FM Hosting, LLM Inference

The AI Engineer 5 (FM Hosting, LLM Inference) joins the Intelligent Foundations and Experiences team to develop and deploy proprietary AI solutions. You will design, test, and support software components including foundation model training, large language model inference, agents, multi-agent workflows, similarity search, and guardrails. Key responsibilities involve inventing optimization techniques to improve performance, scalability, and cost for production systems, while leading cost-performance governance reviews and mentoring senior engineers. You will utilize a technical stack including AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, Python, Go, Scala, CUDA, and Java. The role focuses on building scalable, high-performance AI infrastructure and multi-model orchestration pipelines to solve complex technical problems, integrating LLMs and domain-specific models into unified systems to enhance products and customer interactions.

What does a AI Engineer earn in New York?

Median $245450 from 61 postings across 10 companies.

See salary data

What you'll do

  • Design, develop, test, deploy, and support AI software components including foundation model training and LLM inference.
  • Implement multi-model orchestration pipelines integrating LLMs, vector search, and domain-specific models into unified systems.
  • Develop state-of-the-art optimization techniques to improve scalability, cost, latency, and throughput of production AI systems.
  • Establish and lead cost-performance governance reviews to track GPU utilization and inference cost efficiency.
  • Lead design councils and review boards to ensure technical consistency and compliance with AI engineering standards.
  • Mentor Principal and Manager-level AI engineers to foster cross-domain learning and elevate organizational technical maturity.
  • Contribute to the technical vision and long-term roadmap of foundational AI systems.

What we're looking for

  • Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus 6 years of experience developing AI and ML algorithms.
  • Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus 4 years of experience developing AI and ML algorithms.
  • At least 6 years of experience programming with Python, Go, Scala, CUDA, or Java.
  • Experience leading development of AI systems with tradeoff decisions around cost, latency, throughput, and accuracy (preferred).
  • 7+ years of experience deploying scalable and responsible AI solutions on cloud platforms (preferred).
  • Experience designing, developing, delivering, and supporting complex AI systems (preferred).
  • Experience developing AI and ML algorithms using Python, C++, C#, Java, CUDA, or Golang (preferred).
  • Experience developing and applying state-of-the-art techniques for optimizing training and inference software (preferred).

More like this

Similar roles

AI Engineer 5, FM Hosting, LLM Inference

Capital One Financial

McLean, VA +3 17 days ago $229,900–$262,400
Python Go Scala CUDA Java C++ C# PyTorch Huggingface AWS VectorDBs LLM Inference Similarity Search Multi-agent workflows Model Optimization Model Evaluation Observability GPU Utilization
6+ yrs exp

AI Engineer 5

Capital One Financial

San Jose, CA +3 23 days ago $229,900–$262,400
Python Go Scala CUDA Java C++ C# PyTorch AWS Huggingface VectorDBs LLM Inference Agentic AI Similarity Search Model Optimization GPU Utilization
6+ yrs exp

AI Engineer 5

Capital One Financial

Cambridge, MA +3 23 days ago $229,900–$262,400
Python Go Scala CUDA Java C++ C# PyTorch Huggingface AWS VectorDBs LLM Inference Agentic AI Similarity Search Model Optimization Model Evaluation Observability GPU Utilization
6+ yrs exp

AI Engineer 5

Capital One Financial

McLean, VA +3 20 days ago $229,900–$262,400
Python Go Scala CUDA Java C++ C# PyTorch Huggingface AWS VectorDBs LLM Inference Agentic AI Similarity Search Model Optimization GPU Utilization
6+ yrs exp

AI Engineer 5

Capital One Financial

New York, NY +4 23 days ago $229,900–$262,400
Python Go Scala CUDA Java C++ C# PyTorch Huggingface AWS VectorDBs LLM Reinforcement Learning Model Optimization Similarity Search Multi-agent Workflows GPU Utilization
6+ yrs exp

AI Engineer 5, FM Hosting, LLM Inference

Capital One Financial

San Jose, CA +3 23 days ago
LLM Inference Python PyTorch CUDA VectorDBs AWS Huggingface Go Scala Java C++ C# Model Optimization Similarity Search
6+ yrs exp