Senior Software Engineer, Machine Learning Platform

Chime

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
San Francisco, CA
Posted
6 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $208k
$158k most similar roles pay here $274k

This listing doesn't post a salary. Most similar roles pay $169,350–$246,150.

Based on 240 similar postings.

Employer

About Chime

Chime is a financial technology company offering mobile-first banking services including fee-free checking accounts, savings accounts, and a secured credit builder card through partner banks. Industry: Financial Technology & Neobanking

Chime currently has 36 open roles on FindRole.

Most-posted roles

View all roles at Chime

At a glance

TL;DR · Senior Software Engineer, Machine Learning Platform

As a Senior Software Engineer on the Machine Learning Platform team, you will design and build scalable infrastructure to power machine learning across the company. You will develop systems for model training, feature computation, real-time inference, and agentic orchestration while ensuring high standards for observability, governance, and cost efficiency. Your daily work involves building distributed training and large-scale processing systems using Ray or Spark, managing infrastructure as code with Terraform, and developing data ingestion pipelines using Kinesis, Kafka, or Flink. You will also build evaluation frameworks for non-deterministic AI systems and improve CI/CD workflows. The role requires proficiency in Python, Go, Scala, or Java, along with experience in Docker, Kubernetes, and AWS. You will solve complex problems involving LLM application patterns, including retrieval-augmented generation and multi-step agentic workflows, to enhance the internal developer experience.

What does a Software Engineer earn in California?

Median $214000 from 778 postings across 63 companies.

See salary data

What you'll do

  • Design and operate scalable machine learning and AI infrastructure on AWS.
  • Build shared platform capabilities for LLM and agentic workloads including model access and workflow orchestration.
  • Develop evaluation frameworks for non-deterministic AI systems using offline benchmarks and human feedback.
  • Establish observability, reliability, and governance for models covering latency, cost, and safety.
  • Build distributed training, batch inference, and large-scale processing systems using Ray or Spark.
  • Maintain infrastructure as code using Terraform to support platform components.
  • Develop data ingestion and streaming systems using technologies like Kinesis, Kafka, or Flink.
  • Improve CI/CD workflows for machine learning models and AI applications.

What we're looking for

  • 5+ years of experience in ML or AI infrastructure, platform engineering, distributed systems, or production ML systems.
  • Experience designing distributed systems and large-scale data or compute platforms on AWS using Spark or Ray.
  • Knowledge of the machine learning development lifecycle including preprocessing, training, evaluation, deployment, and monitoring.
  • Working knowledge of LLM application patterns such as RAG, structured outputs, tool calling, and agent orchestration.
  • Experience designing production systems integrating ML or foundation models through reliable APIs, workflows, and data contracts.
  • Hands-on experience with CI/CD pipelines, DevOps practices, and infrastructure as code using Terraform.
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Strong programming skills in Python, Go, Scala, Java, or similar languages.
  • Experience shipping LLM-powered or agentic systems to production (preferred).
  • Experience with model gateways, prompt lifecycle management, retrieval/vector search, and agent orchestration frameworks (preferred).
  • Experience building evaluation, tracing, and observability for non-deterministic AI systems (preferred).
  • Familiarity with managed or self-hosted foundation model infrastructure like Amazon Bedrock or SageMaker (preferred).
  • Experience operating GPU-based workloads and optimizing training/inference performance; CUDA experience is a plus (preferred).

More like this

Similar roles

Senior Machine Learning Engineer, AI Platform

Smartly

Helsinki, Uusimaa, Finland 98 days ago
Machine Learning Python PyTorch TensorFlow MLOps AWS GCP Generative AI Computer Vision Natural Language Processing C++ Java MLflow Kubeflow Data Pipelines Linear Algebra Statistics Calculus
5+ yrs exp Hybrid

Senior Machine Learning Engineer

Apple Inc

Seattle, WA 7 days ago $175,000$308,500
Python PyTorch JAX TensorFlow Go Rust Spark Daft Polars DuckDB Kubernetes Docker Parquet Iceberg Delta Lance Ray Data NVIDIA DALI WebDataset Mosaic StreamingDataset Arrow Retrieval-Augmented Generation

Senior Machine Learning Engineer

Cloudflare, Inc

Austin, TX 77 days ago
Python PyTorch TensorFlow Scikit-Learn LangChain LangGraph Autogen LLMs RAG Vector Databases Docker Kubernetes Terraform Cloudflare Workers FastAPI TypeScript React BigQuery Airflow Argo Workflows CI/CD Pytest
3+ yrs exp Hybrid

AWS Agentic AI Platform Engineer

Booz Allen Hamilton

San Antonio, TX +1 27 days ago $99,000$225,000
Python AWS Amazon Bedrock Amazon SageMaker LangChain LlamaIndex LangGraph Docker Kubernetes Terraform CI/CD MLOps LLMOps FastAPI Pydantic OpenSearch Vector Search Retrieval-Augmented Generation Knowledge Graphs Infrastructure as Code
2+ yrs exp

Senior Machine Learning Engineer, AI Platform

Adobe

San Jose, CA 14 days ago $211,800$306,625
Python Go C++ Rust Java Kubernetes Distributed Systems GPU PyTorch FSDP DeepSpeed vLLM TensorRT-LLM Triton Ray Serve Cloud Infrastructure
7+ yrs exp