AI Research Engineer

Datadog

Confirmed live 2 days ago Trusted
Hybrid

Quick summary

Work type
Hybrid
Location
CanadaFrance
Posted
101 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

How this pay compares to similar roles

Similar $179k
$120k most similar roles pay here $236k

This listing doesn't post a salary. Most similar roles pay $142,075–$215,712.

Based on 240 similar postings.

Employer

About Datadog

Datadog, Inc. is an American company that provides an observability service for cloud-scale applications, providing monitoring of servers, databases, tools, and services, through a SaaS-based data analytics platform.

Datadog currently has 266 open roles on FindRole.

Listed pay typically runs $161,000–$205,000 across 138 roles with salary data.

Most-posted roles

View all roles at Datadog

At a glance

TL;DR · AI Research Engineer

As an AI Research Engineer within the Datadog AI Research team, you will partner with Research Scientists to transform research ideas into functional systems by building data pipelines, training infrastructure, and evaluation tools. You will develop multimodal models for cloud observability and train autonomous agents for SRE incident response. Your daily work involves implementing models, running large-scale experiments, and creating simulation environments for agent training. The role requires proficiency in Python and a systems language like Rust, C++, or Go, alongside experience with PyTorch or JAX. You will utilize tools such as Ray, Slurm, Megatron-LM, DeepSpeed, and SkyRL to manage distributed training and RL loops. This position addresses complex problems in cloud observability and security by building robust infrastructure for world models and automated systems that improve reliability and performance across distributed environments.

What you'll do

  • Build and operate multimodal data pipelines, training infrastructure, and internal tooling for research models.
  • Implement models and execute large-scale experiments while profiling for reliability, performance, and cost.
  • Develop simulation environments and replay infrastructure to support agent training and evaluation.
  • Orchestrate distributed training and reinforcement learning using frameworks like Ray.
  • Establish automated benchmarks and regression tests for world model predictions and agent performance.
  • Transition research prototypes into hardened, reliable services integrated into Datadog products.
  • Contribute to research publications at top-tier conferences and produce high-quality open-source artifacts.

What we're looking for

  • Proficiency in Python and familiarity with systems languages like Rust, C++, or Go.
  • Experience with distributed computing, RL infrastructure, and ML systems for training and inference at scale.
  • Practical experience operating ML training and inference systems using PyTorch or JAX.
  • Expertise in large-scale model training and fine-tuning using frameworks like Megatron-LM, DeepSpeed, SkyRL, VeRL, or TorchTitan.
  • Experience with techniques such as SFT, RLVR, RLHF, and efficient inference methods like quantization and speculative decoding.
  • Familiarity with Ray, Slurm, or similar frameworks for orchestrating distributed training and RL.
  • Experience building multimodal data pipelines, simulation environments, and evaluation infrastructure.
  • Experience supporting or contributing to research publications at top-tier conferences.

More like this

Similar roles

AI Research Scientist

Datadog

Canada +1 101 days ago
Generative AI Machine Learning PyTorch DeepSpeed Megatron-LM CUDA Reinforcement Learning Distributed Training Foundation Models World Models Data Pipelines GPU Programming SRE Cloud Observability Anomaly Detection Root Cause Analysis
Hybrid

AI Research Scientist

Datadog

Remote (Canada) +2 101 days ago $320,000$400,000
Generative AI Machine Learning Reinforcement Learning PyTorch DeepSpeed Megatron-LM CUDA Distributed Training Foundation Models World Models Multimodal Learning Data Pipelines GPU Programming
Remote

ML Systems Research Engineer, RL / Inference / Agent Systems

Amd

Santa Clara, CA 45 days ago $204,000$306,000
Python PyTorch JAX TensorFlow Reinforcement Learning RLHF LLM Agents Kubernetes Ray Slurm CUDA ROCm HIP Distributed Systems Data Pipelines Model Serving Compiler Optimization Profiling Hardware Engineering
Hybrid

Senior Forward Deployed AI Engineer

Amd

Santa Clara, CA 51 days ago $204,000$306,000
Python C++ Rust C TypeScript CUDA HIP LLM Reinforcement Learning RLHF PPO DPO GRPO PyTorch Hugging Face JAX TensorFlow Ray vLLM Distributed Systems
Hybrid

Senior Research Engineer, Autonomous Vehicles

Nvidia

Santa Clara, CA 37 days ago $184,000$287,500
PyTorch JAX TensorFlow Python C++ CUDA Kubernetes SLURM Deep Learning Reinforcement Learning LLMs MLOps HPC Distributed Training Sim-to-Real Natural Language Processing Graphics
10+ yrs exp

Senior Software Engineer, RL Post-Training Frameworks

Nvidia

Remote (Santa Clara, CA) 21 days ago $184,000$287,500
Reinforcement Learning Python C/C++ PyTorch Kubernetes Ray vLLM SGLang TensorRT-LLM DeepSpeed Megatron-LM NCCL InfiniBand FSDP Distributed Systems High-Performance Computing
5+ yrs exp Remote