Technical Director, Large-Scale AI Model Inferencing

Samsung Semiconductor

Confirmed live today High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CA
Salary
$219,000–$351,000 / yr
Posted
29 days ago
Freshness
Confirmed live today

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $245k
This role $285k
$165k most similar roles pay here $371k

This role pays more than 74% of similar roles. Most pay $205,000–$285,625 — the shaded band above. At the midpoint, this role pays about $285k versus about $245k for comparable roles.

Based on 240 similar postings.

Employer

About Samsung Semiconductor

Samsung Semiconductor is the global semiconductor business unit of Samsung Electronics, designing and manufacturing memory chips, logic semiconductors, and foundry solutions for a broad range of applications.

Samsung Semiconductor currently has 37 open roles on FindRole.

Listed pay typically runs $163,000–$253,000 across 37 roles with salary data.

Most-posted roles

View all roles at Samsung Semiconductor

At a glance

TL;DR · Technical Director, Large-Scale AI Model Inferencing

Technical Director, Large-Scale AI Model Inferencing will join the Memory Solutions Lab and Data Fabric Solutions team as a Principal Engineer to lead the development of full-stack AI memory solutions. This role focuses on solving the challenge of managing massive model states across tiered systems including GPU HBM, host DRAM, CXL-attached pools, and NVMe storage. You will serve as the technical authority connecting complex model architectures—such as Dense Transformers, Mixture-of-Experts, and State Space Models—to memory-system designs. Key responsibilities include defining requirements for inference performance, managing large-scale inference stacks like vLLM and SGLang, and developing tiered memory policies to optimize throughput and cost. Required expertise includes deep knowledge of KV-cache management, CUDA Graphs, PyTorch Profiler, and hardware-software co-design. You will translate model behavior into product roadmaps while mentoring engineers on the boundary between AI models and memory systems.

What you'll do

  • Translate AI model architectures into specific memory product requirements and internal reference designs.
  • Model the memory footprint, bandwidth demands, and access patterns for frontier open-weight models.
  • Develop tiered memory systems using HBM, host DRAM, CXL-attached pools, and NVMe/SSD storage.
  • Design expert-weight offloading solutions and cache policies for Mixture-of-Experts (MoE) serving.
  • Drive performance engineering for inference stacks including vLLM, SGLang, and TensorRT-LLM.
  • Create analytical models to predict system behavior and TCO improvements before hardware is finalized.
  • Establish multi-year technical strategies and lead architecture reviews for AI memory solutions.
  • Serve as the primary technical authority in high-level discussions with customers and partners.

What we're looking for

  • BS in Computer, Electrical, Electronic Engineering, or Computer Science with 10+ years of relevant experience.
  • MS in Computer, Electrical, Electronic Engineering, or Computer Science with 8 years of relevant experience (preferred).
  • 12+ years of experience in systems engineering.
  • 4+ years of hands-on experience in large-scale LLM inference or GPU systems performance.
  • First-principles understanding of transformer-class model internals including KV-cache math and MoE behavior.
  • Code-level expertise in the memory-management internals of major inference stacks like vLLM, SGLang, TensorRT-LLM, or llama.cpp.
  • Proven track record building systems software at the memory, storage, or I/O layers such as caches, tiering, and paging.
  • Experience with State Space Models, CXL memory pooling, or NVMe/SSD as a KV cache tier (preferred).

More like this

Similar roles

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +5 64 days ago $184,000–$287,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Director, AI/ML Engineering

Carmax

Richmond, VA 14 days ago $181,400–$290,200
Generative AI Agentic AI Machine Learning LLMs ML Ops LangChain AutoGen OpenAI Anthropic AWS Bedrock Azure OpenAI Computer Vision Agile Data Governance Feature Engineering
10+ yrs exp Hybrid

Director, AI/ML Engineering

Fidelity Financial Services

Westlake, TX 15 days ago
AI Machine Learning MLOps Python AWS SageMaker Lambda Glue Step Functions S3 EMR Athena Snowflake SQL PyTorch Amazon Bedrock LLMs RAG Vector Databases Tableau Kinesis Docker Jenkins Artifactory SonarQube GitHub CI/CD
6+ yrs exp

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 72 days ago $224,000–$356,500
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Director, AI Engineering

Humana

Louisville, KY +2 16 days ago
Generative AI LLM Orchestration Python TensorFlow PyTorch AWS Azure GCP Retrieval-Augmented Generation Fine-tuning CI/CD Automated Testing Observability Machine Learning Data Engineering DevOps
10+ yrs exp Hybrid