ML Framework Engineer

Apple Inc

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Cupertino, CA
Salary
$150,400–$225,300 / yr
Posted
17 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $226k
This role $188k
$135k most similar roles pay here $293k

This role pays less than 80% of similar roles. Most pay $196,562–$254,750 — the shaded band above. At the midpoint, this role pays about $188k versus about $226k for comparable roles.

Based on 240 similar postings.

Employer

About Apple Inc

Apple Inc. is a multinational technology company known for designing and manufacturing consumer electronics, software, and online services, including the iPhone, Mac, iPad, and App Store. Industry: Consumer Electronics & Software

Apple Inc currently has 1984 open roles on FindRole.

Listed pay typically runs $175,000–$277,600 across 1590 roles with salary data.

Most-posted roles

View all roles at Apple Inc

At a glance

TL;DR · ML Framework Engineer

The ML Framework (MetalLM) Engineer joins the GPU, Graphics and Machine Learning team to enable high-performance, distributed inference of GenAI applications like LLMs on Private Cloud Compute. You will develop kernel and compiler level optimizations while implementing features for Metal device backends to accelerate training frameworks such as PyTorch and JAX. Day-to-day responsibilities include optimizing code for scalable inference using data, tensor, pipeline, and expert parallelism, alongside applying techniques like quantization, compression, and speculation to improve throughput and reduce latency. The role requires expertise in C/C++/ObjC, GPU kernel development with Metal or CUDA, and system level programming. You will collaborate with hardware and compiler teams to align software performance with Apple Silicon capabilities. Preferred skills include experience with Triton, OpenXLA, LLVM/MLIR, and contributions to major AI frameworks like PyTorch, JAX, or TensorFlow.

What you'll do

  • Optimize code for scalable ML inference using distributed compute strategies like data, tensor, pipeline, and expert parallelism.
  • Develop kernel and compiler level optimizations to ensure peak performance across various server hardware families.
  • Apply model optimization techniques including speculation, quantization, and compression to maximize throughput and minimize latency.
  • Analyze and improve key performance metrics such as end-to-end latency, TTFT, TBOT, memory footprint, and compute efficiency.
  • Implement features for the Metal device backend to accelerate machine learning training technologies.
  • Align software performance with hardware capabilities through close collaboration with hardware, compiler, and systems teams.

What we're looking for

  • At least 2 years of programming and problem-solving experience with C, C++, or Objective-C.
  • Experience with GPU kernel development and optimizations using compute programming models like Metal or CUDA.
  • Experience with system-level programming and computer architecture.
  • Experience with distributed training or inference techniques.
  • Experience with graph compilers such as Triton, OpenXLA, or LLVM/MLIR (preferred).
  • Contributions to an AI framework such as PyTorch, JAX, or TensorFlow (preferred).
  • Good understanding of machine learning fundamentals (preferred).

More like this

Similar roles

ML Software Engineer

Apple Inc

Seattle, WA 42 days ago $142,300$263,300
Swift C++ Python Go Rust Java gRPC Protocol Buffers OpenTelemetry Splunk Distributed Systems ML Inference Quantization GPU Acceleration Swift Concurrency XPC Instruments Concurrency Multi-node Clusters
2+ yrs exp

Staff ML Engineer, Ads ML Infrastructure

Apple Inc

Cupertino, CA 66 days ago $184,700$324,800
Machine Learning ML Serving ONNX Runtime TensorRT vLLM Ray GPU Kernels Distributed Systems Feature Stores Federated Learning Privacy-Preserving ML RPC
8+ yrs exp