Site Reliability Engineer, AI Platform & Cloud

Morgan Stanley

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Alpharetta, GA
Posted
143 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $185k
$141k most similar roles pay here $238k

This listing doesn't post a salary. Most similar roles pay $150,368–$219,125.

Based on 238 similar postings.

Employer

About Morgan Stanley

Morgan Stanley is a global financial services firm providing investment banking, securities, wealth management, and investment management services to corporations, governments, institutions, and individuals. Industry: Investment Banking & Financial Services

Morgan Stanley currently has 38 open roles on FindRole.

Listed pay typically runs $150,000–$210,000 across 31 roles with salary data.

Most-posted roles

View all roles at Morgan Stanley

At a glance

TL;DR · Site Reliability Engineer, AI Platform & Cloud

Site Reliability Engineer (SRE) - AI Platform & Cloud is a Director level position within the Technology division. You will join the AI Platform team to support, scale, and harden infrastructure powering AI/ML systems in a regulated financial environment. Your daily responsibilities include operating and maintaining infrastructure for GenAI applications, building automation to reduce manual toil, and managing infrastructure-as-code for compute, storage, and GPU clusters. You will establish SLOs, lead incident responses, perform capacity planning, and optimize cost versus performance tradeoffs. The role requires expertise in Kubernetes, Docker, Terraform, Helm, and Ansible, alongside proficiency in Python, Go, or Java. You will work with tools like Prometheus, Grafana, and the ELK stack while managing data pipelines using Kafka or Spark to ensure reliable production workloads for training, inference, and model serving.

What does a Site Reliability Engineer earn?

Median $186200 from 131 postings across 36 companies.

See salary data

What you'll do

  • Operate, monitor, and maintain infrastructure supporting GenAI applications including training, inference, and data ingestion.
  • Design and build automation to reduce manual toil across core platform capabilities.
  • Develop and maintain Infrastructure-as-Code (IaC) for managing compute, storage, network, and GPU clusters.
  • Establish and enforce SLOs, SLIs, SLAs, error budgets, and automated alerting dashboards.
  • Lead incident response, root cause analysis, postmortems, and systemic remediation efforts.
  • Perform capacity planning, workload scheduling, and resource forecasting to optimize cost versus performance.
  • Harden systems for security, compliance, auditability, and data governance in a regulated environment.
  • Define disaster recovery strategies, including backup/restore practices and fault tolerance mechanisms.

What we're looking for

  • Bachelor’s or Master’s degree in Computer Science or a related field, or equivalent job experience.
  • 5 years of production experience in SRE, infrastructure, or operations for large-scale systems.
  • Strong programming and scripting skills in Python, Go, Java, or equivalent languages.
  • Deep experience with containerization using Docker and orchestration via Kubernetes.
  • Proficiency in Infrastructure-as-Code (IaC) tools such as Terraform, Helm, CloudFormation, or Ansible.
  • Experience with monitoring, observability, and logging tools like Prometheus, Grafana, ELK/EFK, or Datadog.
  • Knowledge of networking and systems engineering including TCP/IP, DNS, routing, and load balancing.
  • Experience in regulated environments such as financial services, compliance, audit, and security.

More like this

Similar roles

Site Reliability Engineer

Morgan Stanley

Alpharetta, GA 22 days ago
Python Shell Scripting Perl Ruby Java C# AWS Azure Jenkins Splunk DB2 Oracle Sybase Autosys Linux Unix Windows Agile Scrum Web Services MQ
5+ yrs exp

Site Reliability Engineering Lead

US Bank

Atlanta, GA +2 3 days ago $111,605$131,300
SRE DevOps AWS Azure Kubernetes Docker Terraform Ansible Python PowerShell Shell Scripting CI/CD GitHub Actions Azure DevOps Jenkins GitLab Datadog Splunk Dynatrace Grafana Prometheus CloudWatch Azure Monitor OpenTelemetry ServiceNow Jira REST APIs SQL
6+ yrs exp

Site Reliability Engineer

CME Group

Chicago, IL 7 days ago $103,500$172,500
GCP AWS Azure Kubernetes Docker Python Java Linux Unix Oracle BigQuery OpenTelemetry Splunk Prometheus Grafana UC4 Automic Bamboo JIRA Git CI/CD

Site Reliability Engineer

Berkeley Research Group

Remote 77 days ago $130,000$160,000
Azure Kubernetes CI/CD GitHub Actions GitLab CI Golang Ruby Python AWS GCP Infrastructure as Code Datadog OpsGenie PagerDuty SRE Incident Management
5+ yrs exp Remote

Site Reliability Engineer

Balyasny Asset Management

Warsaw, Poland 78 days ago
Prometheus Grafana Loki Tempo OTEL Kubernetes Docker AWS Python Bash Go CI/CD DevOps SRE Agile
5+ yrs exp