Browse tech roles

Basic role filtering by workplace, salary floor, and post age. For full AI matching and advanced filtering upload your resume using AI Match.

5 of up to 20 (filtered)

Engineering Manager, LLM Inference & Deployment at Scale

Nvidia

Santa Clara, CA 14 days ago $224,000$356,500
Actively hiring Confirmed live yesterday High trust Above market
LLMs VLMs TensorRT TensorRT-LLM vLLM SGLang Quantization Speculative Decoding Continuous Batching Prefix Caching KV-cache Optimization Distributed Computing GPU Cluster Orchestration Model Serving Inference Optimization Deep Learning
8+ yrs exp Hybrid

Software Engineer, Inference (AI Data Engineering)

SpaceX

Palo Alto, CA 23 days ago $135,000$175,000
Actively hiring Confirmed live 2 days ago High trust Competitive pay
SGLang vLLM TensorRT-LLM Triton Rust C++ Python Go gRPC REST Docker Kubernetes PostgreSQL ClickHouse MongoDB CI/CD GPU Kernels Quantization Speculative Decoding
2+ yrs exp

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA +4 37 days ago $184,000$287,500
Actively hiring Confirmed live 2 days ago High trust Competitive pay
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Communication Deep Learning Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Deep Learning Inference

Nvidia

Santa Clara, CA 45 days ago $224,000$356,500
Actively hiring Confirmed live yesterday High trust Above market
CUDA Triton CUTLASS C++ Python vLLM SGLang FlashInfer TensorRT-LLM PyTorch NCCL NVSHMEM NIXL Multi-GPU Distributed Inference Agile
6+ yrs exp Hybrid

Engineering Manager, Inference Benchmarking

Nvidia

Remote (Santa Clara, CA) +1 106 days ago $224,000$356,500
Actively hiring Confirmed live yesterday High trust Above market
LLM vLLM TRT-LLM SGLang Kubernetes Prometheus ZMQ DCGM PyNVML Helm Distributed Systems Microservices open-source Inference Infrastructure Benchmarking Computer Vision
8+ yrs exp Remote