Senior Staff Engineer, ML Inference

Shopify

Confirmed live 2 days ago High trust
Remote

Quick summary

Work type
Remote
Location
Remote
Posted
134 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

How this pay compares to similar roles

Similar $236k
$171k most similar roles pay here $314k

This listing doesn't post a salary. Most similar roles pay $212,700–$259,212.

Based on 240 similar postings.

Employer

About Shopify

Shopify is a leading global commerce platform that enables businesses of all sizes to start, grow, and manage their retail operations online and in-person. It provides tools for storefronts, payments, shipping, and marketing to millions of merchants worldwide.

Shopify currently has 19 open roles on FindRole.

Most-posted roles

View all roles at Shopify

At a glance

TL;DR · Senior Staff Engineer, ML Inference

As a Senior Staff Engineer - ML Inference, you will join a remote-first team to architect, optimize, and own high-performance production machine learning inference systems. You will build the engine for real-time AI systems by designing for high throughput, ultra-low latency, and global reliability while driving significant cost and performance optimizations. Your daily work involves implementing advanced techniques like pruning, quantization, distillation, and batching to serve models in production. To achieve these goals, you will utilize technologies including CUDA, TensorRT, Triton, TVM, and custom GPU kernels, while writing code in Python or C++. You will collaborate with infrastructure and product teams to deploy and scale large-scale ML workloads. This role focuses on the technical challenge of optimizing inference across diverse hardware and managing high-volume queries within a complex commerce environment.

What does a Engineer earn in Remote?

Median $189520 from 49 postings across 15 companies.

See salary data

What you'll do

  • Architect and own high-performance production ML inference systems for high throughput and low latency.
  • Optimize model performance using CUDA, TensorRT, Triton, TVM, and custom GPU kernels.
  • Implement advanced techniques like pruning, quantization, distillation, and batching to improve serving efficiency.
  • Drive cost optimization and system efficiency to reduce cloud spend and carbon footprint.
  • Lead deep performance investigations to resolve bottlenecks in large-scale ML workloads.
  • Define the technical strategy and culture for ML inference across the organization.
  • Partner with cross-functional teams to deploy, benchmark, and scale cutting-edge models.

What we're looking for

  • Proven expertise in building and optimizing large-scale ML inference systems with measurable performance wins.
  • Experience in production model serving, runtime optimization, and acceleration using GPUs (CUDA, TensorRT).
  • Strong software engineering skills in Python, C++, or other relevant languages with a distributed computing mindset.
  • Demonstrated leadership in architecting and scaling reliable, real-time inference handling millions of queries daily.
  • Advanced understanding of model compression techniques including pruning, quantization, distillation, and batching.
  • Ability to collaborate cross-functionally with ML research, infrastructure, and product teams.
  • Experience optimizing inference across diverse hardware such as NVIDIA, AMD, ARM, or cloud TPUs.
  • Familiarity with MLOps pipelines, monitoring, observability, and auto-scaling for inference platforms.

More like this

Similar roles

Senior Staff Machine Learning Engineer

DoorDash, Inc

San Francisco, CA +1 100 days ago $242,800$357,000
Large Language Models Python PyTorch TensorFlow XGBoost Natural Language Processing Information Retrieval Ranking and Relevance Recommendation Systems RAG Prompt Engineering C++ Java A/B Testing Feature Engineering Deep Learning Sequence Modeling
5+ yrs exp

Staff ML Engineer, Ads ML Infrastructure

Apple Inc

New York, NY 66 days ago $184,700$324,800
Machine Learning ML Serving ONNX Runtime TensorRT vLLM Ray GPU Kernels Distributed Systems Feature Stores Federated Learning Privacy-Preserving ML RPC Agile
8+ yrs exp

Staff ML Engineer, Ads ML Infrastructure

Apple Inc

Cupertino, CA 66 days ago $184,700$324,800
Machine Learning ML Serving ONNX Runtime TensorRT vLLM Ray GPU Kernels Distributed Systems Feature Stores Federated Learning Privacy-Preserving ML RPC
8+ yrs exp

Senior Staff ML Engineer

GEICO

Remote (Palo Alto, CA) 45 days ago $150,000$300,000
Generative AI LLM Agentic Workflows Python Java RAG LangSmith LangGraph Kubernetes CI/CD AWS Azure Prompt Engineering
8+ yrs exp Remote

Senior Staff Machine Learning Engineer

GEICO

Palo Alto, CA +1 98 days ago $150,000$300,000
Generative AI LLMs Python Java C++ C# AWS Azure Kubernetes CICD Kafka Spark Ray Airflow Temporal elastic search Qdrant Snowflake PostgreSQL MongoDB Cassandra
10+ yrs exp

Senior Staff Machine Learning Engineer

GEICO

Palo Alto, CA 135 days ago $150,000$300,000
Generative AI LLMs Python Java C++ C# AWS Azure Kubernetes CICD Kafka Spark Ray Airflow Temporal PostgreSQL MongoDB Cassandra Snowflake elastic search Qdrant
10+ yrs exp