Senior Deep Learning Engineer, Cosmo 3D Spatial

Nvidia

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Santa Clara, CA
Salary
$224,000–$356,500 / yr
Employment
Full-time
Posted
18 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $229k
This role $290k
$153k most similar roles pay here $378k

This role pays more than 95% of similar roles. Most pay $202,475–$254,750 — the shaded band above. At the midpoint, this role pays about $290k versus about $229k for comparable roles.

Based on 240 similar postings.

Employer

About Nvidia

Nvidia is a leading designer of graphics processing units (GPUs) and system-on-chip units, powering gaming, professional visualization, data centers, and artificial intelligence workloads. Industry: Semiconductors & AI Computing

Nvidia currently has 1150 open roles on FindRole.

Listed pay typically runs $184,000–$287,500 across 912 roles with salary data.

Most-posted roles

View all roles at Nvidia

At a glance

TL;DR · Senior Deep Learning Engineer, Cosmo 3D Spatial

As a Senior Deep Learning Engineer on the Cosmo Engineering team, you will own the 3D data engine and end-to-end training verification for models providing 3D spatial reasoning and perception capabilities. You will build annotation and auto-labeling pipelines to produce 3D-grounded supervision, manage large-scale vision-language training data, and develop evaluation suites including benchmarks like CV-Bench and RoboSpatial. The role involves translating research hypotheses into experiments, managing multi-node GPU clusters, and optimizing data throughput for pre-training and supervised fine-tuning. You will utilize Python with PyTorch or JAX on Linux to solve problems in 3D computer vision, multi-view geometry, and structure-from-motion. Key technical requirements include experience with distributed training strategies like FSDP, as well as expertise in point cloud processing, depth estimation, and building large-scale multimodal data pipelines for robots and autonomous vehicles.

What you'll do

  • Source, curate, and filter large-scale real-world image and video corpora into vision-language training data for 3D spatial reasoning.
  • Build automated annotation and labeling pipelines to produce 3D-grounded supervision and multi-modal training data at scale.
  • Manage end-to-end data quality through semantic deduplication, automated scoring, and balanced sampling strategies for pre-training and fine-tuning.
  • Validate full model training pipelines by monitoring performance, diagnosing anomalies, and attributing capability changes to specific data decisions.
  • Build and operate 3D spatial evaluation suites and public benchmarks to measure model reasoning and geometry capabilities.
  • Translate research hypotheses from scientists into large-scale dataset experiments and measurable technical requirements for the Cosmos platform.
  • Optimize data throughput and sharding on multi-node GPU clusters to ensure high performance during long training runs.

What we're looking for

  • MS or PhD in Computer Science, Electrical/Computer Engineering, Robotics, or a related field, or equivalent experience.
  • 12+ years of proven experience building deep learning systems in Python with PyTorch or JAX on Linux.
  • Deep expertise in 3D computer vision, multi-view geometry, structure-from-motion, SLAM, depth/camera pose estimation, point cloud processing, or 3D reconstruction.
  • Hands-on experience with vision-language models and building training data to improve visual grounding and reasoning quality.
  • Demonstrated experience building large-scale multimodal data pipelines including distributed video/image processing, deduplication, and automated quality metrics.
  • Experience running and validating large model training on multi-GPU, multi-node clusters using distributed training and sharding strategies.
  • Excellent written and verbal communication skills to partner with research scientists and translate research into engineering execution.
  • PhD and/or publications at CVPR, ICCV, ECCV, NeurIPS, ICLR, or CoRL in 3D vision, multimodal learning, or embodied AI (preferred).

More like this

Similar roles

Senior Deep Learning Software Engineer, Inference

Nvidia

Remote (CA) +4 37 days ago $152,000–$241,500
CUDA C++ Python PyTorch vLLM SGLang Triton CUTLASS NCCL NVSHMEM Deep Learning LLM Generative AI GPU Programming Performance Optimization Profiling Agile
5+ yrs exp Remote