below marketmost similar roles pay hereabove market
This role pays more than
60%
of similar roles. Most pay
$202,800–$254,750
— the blue band above.
At the midpoint, this role pays about
$234k
versus about
$229k
for comparable roles.
Based on 240 similar postings.
Employer
About Capital One Financial
Capital One Financial is a bank holding company specializing in credit cards, auto loans, banking, and savings products, known for its data-driven approach to consumer and commercial finance. Industry: Financial Services & Banking
Capital One Financial currently has
1917 open roles
on FindRole.
Listed pay typically runs
$197,300–$225,100
across 1045 roles with salary data.
The Applied Researcher 4 (AI Foundations: Multimodal Guardrails and VLM's) joins the AI Foundations team to develop trustworthy and reliable AI systems. This role involves partnering with cross-functional teams of data scientists, software engineers, and product managers to build AI foundation models through all phases of development, including design, training, evaluation, validation, and implementation. You will leverage a technology stack including Pytorch, AWS Ultraclusters, Huggingface, and Lightning to extract insights from large volumes of numeric and textual data. Key responsibilities include conducting high-impact applied research, translating complex technical work into business goals, and preparing internal tech reports or conference submissions. The work focuses on pushing state-of-the-art AI developments into production to improve customer experiences and solve complex problems involving large deep learning models.
Build AI foundation models through all phases of development, including design, training, evaluation, and implementation.
Leverage Pytorch, AWS Ultraclusters, Huggingface, and Lightning to extract insights from large numeric and textual datasets.
Conduct high-impact applied research to integrate the latest AI developments into next-generation customer experiences.
Develop and deliver AI-powered products and scalable models for production use.
Translate complex technical research into tangible business goals for stakeholders.
Prepare internal technical reports or conference submissions summarizing novel methods and research findings.
Own and pursue a research agenda by identifying impactful problems and autonomously carrying out long-running projects.
What we're looking for
Must have a PhD in Electrical Engineering, Computer Engineering, Computer Science, AI, Mathematics, or a related field.
Must have an M.S. in Electrical Engineering, Computer Engineering, Computer Science, AI, Mathematics, or a related field plus 2 years of experience in Applied Research.
Must have hands-on experience developing AI foundation models and solutions using open-source tools and cloud computing platforms.
Must have a deep understanding of the foundations of AI methodologies.
Must have experience building large deep learning models for language, images, events, or graphs.
Must have expertise in one or more of the following: training optimization, self-supervised learning, robustness, explainability, or RLHF.
Must have a track record of delivering models at scale in terms of training data and inference volumes.
Must have experience delivering libraries, platform-level code, or solution-level code to existing products.
Must have a track record of producing high-quality machine learning ideas, such as first-author publications or projects.
Must possess the ability to own and pursue a research agenda and autonomously carry out long-running projects.
PhD in Computer Science, Machine Learning, Computer Engineering, Applied Mathematics, or Electrical Engineering (preferred).
LLM PhD focus on NLP or Masters with 5 years of industrial NLP research experience (preferred).
Multiple publications on topics related to the pre-training of large language models (preferred).
Member of a team that has trained a large language model from scratch with 10B+ parameters and 500B+ tokens (preferred).
Publications in deep learning theory (preferred).
Publications at ACL, NAACL, EMNLP, Neurips, ICML, or ICLR (preferred).
PhD focused on topics related to optimizing training of very large deep learning models (preferred).
Multiple years of experience and/or publications on Model Sparsification, Quantization, Training Parallelism, Gradient Checkpointing, or Model Compression (preferred).
Deep knowledge of deep learning algorithmic and/or optimizer design (preferred).
Experience with compiler design (preferred).
PhD focused on topics related to guiding LLMs with Supervised Finetuning, Instruction-Tuning, Dialogue-Finetuning, or Parameter Tuning (preferred).
Demonstrated knowledge of principles of transfer learning, model adaptation, and model guidance (preferred).
Experience deploying a fine-tuned large language model (preferred).
Demonstrated ability to reproduce and extend peer-reviewed AI research using modern open-source frameworks (preferred).
Experience designing controlled experiments and documenting reproducibility results for internal or public research (preferred).