Quick summary
- Work type
- Remote
- Location
- Remote
- Salary
- $230,000–$322,000 / yr
- Posted
- 177 days ago
- Freshness
- Confirmed live 2 days ago
Market check
Salary context
How this pay compares to similar roles
This role pays more than 85% of similar roles. Most pay $198,375–$254,750 — the shaded band above. At the midpoint, this role pays about $276k versus about $227k for comparable roles.
Based on 240 similar postings.
Employer
About Reddit
Reddit is a social news aggregation and discussion platform where users share content, vote on posts, and engage in community conversations across thousands of interest-based forums called subreddits.
Reddit currently has 77 open roles on FindRole.
Listed pay typically runs $217,000–$303,400 across 77 roles with salary data.
Most-posted roles
- Software Engineer 19
- Data Scientist 8
- Machine Learning Engineer 5
- Frontend Engineer 4
- Machine Learning Systems Engineer 4
At a glance
TL;DR · Staff Machine Learning Systems Engineer
Staff Machine Learning Systems Engineer As part of the Machine Learning Platform team, the Staff Machine Learning Systems Engineer will lead the development of a platform for large-scale machine learning models. The role involves designing end-to-end MLOps patterns to improve developer velocity, building and supporting a graph ML codebase that abstracts common patterns, and collaborating on performance tuning to optimize training time and GPU costs in distributed environments. Key responsibilities include optimizing batch data processing and architecting pipelines for massive graph data structures. The candidate will utilize technologies including Python, PyTorch, TensorFlow, Ray, Kubernetes, Terraform, and GCP BigQuery. Experience with MLOps tools like MLflow or Wandb is required. Additionally, the role involves solving complex problems in content discovery and recommendation systems by managing large-scale infrastructure for model management, experiment tracking, and distributed training across billions of nodes.
Skills
What you'll do
- Design end-to-end MLOps patterns including data preparation, model management, and experiment tracking.
- Develop and support a graph ML codebase that abstracts common patterns for better scalability.
- Perform performance tuning to improve training time, efficiency, and GPU costs in distributed environments.
- Optimize batch data processing using tools like Apache Beam, Spark, and Ray Data.
- Architect pipelines to build and maintain massive graph data structures with billions of nodes and edges.
- Integrate MLOps tools for experiment tracking, model serving, and model registries.
What we're looking for
- 8+ years of experience in ML infrastructure, including model training and deployment.
- Hands-on experience with ML optimization, specifically memory and GPU profiling.
- Deep experience with cloud technologies like GCP BigQuery, Google Cloud Storage, and Terraform.
- Experience administering MLOps tools for experiment tracking, model serving, and registries like MLflow or Wandb.
- Proficiency in common ML programming languages and frameworks such as Python, PyTorch, and TensorFlow.
- Deep experience with distributed training frameworks including Ray and Kubernetes.
- Ability to optimize batch data processing using tools like Apache Beam, Spark, or Ray Data.
- Experience with graph databases (Neo4j, JanusGraph, TigerGraph) and graph neural networks is a plus.
More like this
Similar roles
Staff Machine Learning Engineer
Intuit
Staff Machine Learning Engineer
Intuit
Staff Machine Learning Engineer
GEICO
Staff Machine Learning Engineer
Adobe
Staff Machine Learning Engineer
Arm Holdings