PhD Research Intern, Efficient Deep Learning

Nvidia

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Santa Clara, CA
Employment
Full-time
Posted
4 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $192k
$130k $267k
below market most similar roles pay here above market

This listing doesn't post a salary. Most similar roles pay $142,937–$241,750.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1463 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 1096 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · PhD Research Intern, Efficient Deep Learning

PhD Research Intern, Efficient Deep Learning - 2027 will join the Deep Learning Efficiency Research team to research, design, and implement novel methods for efficient deep learning. The role focuses on two core areas: efficient diffusion language models and multimodal generative models, and efficient agentic AI with hybrid inference orchestration across cloud and edge. Key responsibilities include developing sampling efficiency, adaptive unmasking, self-speculation, parallel decoding, and resource-aware agent loops. You will work with large language models, diffusion language models, and multimodal vision-language models. Required skills include large-scale model training, data preparation, and model parallelization using tensor and pipeline methods. The work involves post-training model optimization like pruning, quantization, and NAS to solve technical problems in training, finetuning, and hybrid cloud-edge inference orchestration.

What you'll do

  • Research, design, and implement novel methods for efficient diffusion language models and multimodal generative models.
  • Develop sampling efficiency, adaptive unmasking, self-speculation, and parallel decoding techniques.
  • Build training and distillation pipelines for multimodal generation.
  • Design hybrid inference orchestration for agentic AI across cloud and edge environments.
  • Develop routing and scheduling policies for on-device and cloud expert delegation.
  • Create resource-aware agent loops and adaptive routing systems.
  • Publish original research in top-tier machine learning and computer vision venues.
  • Transfer developed technologies to product groups for real-world application.

What we're looking for

  • Currently pursuing a Ph.D. in Computer Science, Engineering, Electrical Engineering, or a related field.
  • Excellent knowledge of the theory and practice of machine learning and deep learning.
  • Experience with large language models, diffusion language models, multimodal/vision-language models, or agentic systems.
  • Hands-on experience with large-scale model training, including data preparation and model parallelization (tensor and pipeline).
  • Outstanding research track record with at least one top-tier conference publication (e.g., ICML, ICLR, NeurIPS, CVPR, ICCV).
  • Excellent communication skills.
  • Parallel programming experience, such as CUDA (preferred).
  • Experience in hybrid cloud–edge inference, orchestration, adaptive routing, pruning, quantization, NAS, or efficient backbones (preferred).

More like this

Similar roles

Research Scientist, Efficient Deep Learning

Nvidia

Santa Clara, CA 122 days ago
Deep Learning Computer Vision Python PyTorch C++ CUDA Large Language Models Vision-Language Models Pruning Quantization NAS Model Parallelization Generative Models Parallel Programming

PhD Research Intern, AI-Aided Engineering

Nvidia

Santa Clara, CA 4 days ago
Python PyTorch Deep Learning Neural Operators Computational Fluid Dynamics CFD CAD NURBS Geometry Processing Vision Transformers Large Language Models Point Clouds Meshes Agentic AI Physical Simulation Numerical Solvers

PhD Research Intern, Fundamental Generative AI

Nvidia

Santa Clara, CA 25 days ago
Generative AI Python PyTorch C++ CUDA Computer Vision NLP 3D Parallel Programming Machine Learning Deep Learning Computer Architecture Multimodal Data AI for Science

PhD Research Intern, AI Accelerator Design and VLSI

Nvidia

Santa Clara, CA 12 days ago
VLSI Design AI HW/SW Co-Design Python PyTorch SystemVerilog C++ High-Level Synthesis (HLS) Machine Learning Quantization Tensor Decomposition Digital VLSI Circuits Hardware Accelerator Architecture Tapeout