AI Inference Engineer

F5 Inc

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
San Jose, CASeattle, WA
Salary
$176,600–$265,000 / yr
Posted
15 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $197k
This role $221k
$136k most similar roles pay here $279k

This role pays more than 72% of similar roles. Most pay $158,400–$235,750 — the shaded band above. At the midpoint, this role pays about $221k versus about $197k for comparable roles.

Based on 240 similar postings.

Employer

About F5 Inc

F5, Inc. is an American technology company specializing in application security, multi-cloud management, online fraud prevention, application delivery networking, application availability and performance, and network security, access, and authorization.

F5 Inc currently has 25 open roles on FindRole.

Listed pay typically runs $152,200–$228,400 across 25 roles with salary data.

Most-posted roles

View all roles at F5 Inc

At a glance

TL;DR · AI Inference Engineer

The AI Inference Engineer joins the team to bridge the gap between high-performance model development and optimized deployment environments. This role focuses on optimizing Large Language Models for inference across diverse settings, from GPU-rich data centers to resource-constrained edge devices. The engineer will build and maintain robust inference engines using vLLM, TGI, and NVIDIA Triton while designing auto-scaling architectures via Kubernetes for real-time and batch pipelines. Key responsibilities include hardware acceleration for NVIDIA GPUs, Apple Silicon, TPUs, and LPUs, alongside establishing observability frameworks to monitor metrics like Time to First Token and memory bandwidth. Required skills include proficiency in Python, C++, Rust, or Golang, along with experience in Docker, cloud platforms like AWS, GCP, and Azure. The role solves the technical challenge of ensuring enterprise-grade reliability and low-latency performance for large-scale AI applications.

What you'll do

  • Build and maintain robust inference engines using tools like vLLM, TGI, and NVIDIA Triton.
  • Optimize Large Language Models for diverse hardware backends including NVIDIA GPUs, Apple Silicon, TPUs, and LPUs.
  • Design and implement auto-scaling architectures for real-time and batch inference pipelines using Kubernetes.
  • Develop observability frameworks to monitor metrics like Time to First Token (TTFT) and tokens per second.
  • Execute performance and load testing suites to identify bottlenecks and ensure reliability during traffic spikes.
  • Implement advanced optimization techniques such as Speculative Decoding and PagedAttention for model deployment.
  • Improve cost efficiency by maximizing hardware utilization and optimizing resource allocation for inference tasks.

What we're looking for

  • Proficiency in programming languages such as Python, C++, Rust, or Golang for high-performance AI workflows.
  • Hands-on experience with inference tools including vLLM, TensorRT, Llama.cpp, and Ollama.
  • Strong familiarity with infrastructure technologies including Docker, Kubernetes, and cloud platforms like AWS, GCP, and Azure.
  • Comprehensive understanding of GPU and AI hardware, including profiling and optimizing for NVIDIA GPUs and TPUs.
  • Experience deploying Large Language Models using advanced techniques like Speculative Decoding or PagedAttention.
  • Experience contributing to open-source inference libraries or developing hardware-level kernels such as CUDA or Triton.
  • Background in MLOps or SRE roles focused on high-performance AI endpoints and reliability during demand surges.
  • Ability to design scalable solutions for high-throughput inference environments optimized for traffic bursts.

More like this

Similar roles

Senior Software Engineer, AI Inference Systems

Nvidia

Santa Clara, CA 136 days ago $184,000$287,500
vLLM SGLang CUDA Python C++ PyTorch Triton MLIR LLVM XLA Docker Kubernetes Slurm NCCL Nsight Systems CI/CD AWS GCP Azure High-Performance Computing
7+ yrs exp Hybrid

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 31 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

Senior AI Inference Platform Engineer

Apple Inc

Seattle, WA 30 days ago $175,000$308,500
Python Go C++ Triton TensorRT-LLM vLLM Kubernetes Prometheus Grafana Splunk CI/CD Data Pipelines Performance Benchmarking Distributed Systems Capacity Planning
7+ yrs exp

Principal GenAI Inference Optimization Engineer

Amd

San Jose, CA 167 days ago $240,000$360,000
GenAI LLM GPU Architecture vLLM SGLang Triton TensorRT-LLM PyTorch JAX TensorFlow Python C++ CUDA HIP Quantization Distributed Systems RDMA Profiling Benchmarking
Hybrid

Software Engineer, Inference (AI Data Engineering)

SpaceX

Palo Alto, CA 23 days ago $135,000$175,000
SGLang vLLM TensorRT-LLM Triton Rust C++ Python Go gRPC REST Docker Kubernetes PostgreSQL ClickHouse MongoDB CI/CD GPU Kernels Quantization Speculative Decoding
2+ yrs exp