Principal Site Reliability Engineer

Ally Financial

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Salary
$110,000–$180,000 / yr
Posted
3 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $184k
This role $145k
$95k most similar roles pay here $250k

This role pays less than 80% of similar roles. Most pay $151,404–$216,250 — the shaded band above. At the midpoint, this role pays about $145k versus about $184k for comparable roles.

Based on 238 similar postings.

Employer

About Ally Financial

Ally Financial is a US-based digital financial services company offering online banking, auto financing, mortgage, and investment products. It is one of the largest all-digital banks in the United States.

Ally Financial currently has 13 open roles on FindRole.

Listed pay typically runs $110,000–$180,000 across 13 roles with salary data.

Most-posted roles

View all roles at Ally Financial

At a glance

TL;DR · Principal Site Reliability Engineer

The Principal Site Reliability Engineer (SRE) joins a team focused on advancing Generative AI and Machine Learning capabilities in reporting while meeting product and technology PMO needs. This role involves designing and implementing highly available, scalable infrastructure systems for mission-critical production services, including automated deployment pipelines, observability platforms, and disaster recovery protocols. The engineer will lead incident response and postmortem processes to resolve complex distributed system failures, manage service level objectives and error budgets, and build automation tools to eliminate toil. Key technical requirements include expertise in AWS with Terraform or Pulumi, proficiency in Python, Go, or Node, and experience with container orchestration like ECS. The role requires skills in monitoring tools such as Dynatrace, Prometheus, and Grafana, alongside deep knowledge of CI/CD pipelines, GitOps workflows, and distributed systems principles.

What does a Site Reliability Engineer earn?

Median $188000 from 134 postings across 37 companies.

See salary data

What you'll do

  • Design and implement highly available, scalable infrastructure systems for mission-critical production services.
  • Build automated deployment pipelines, observability platforms, and disaster recovery systems.
  • Lead incident response and postmortem processes to identify root causes of distributed system failures.
  • Develop and maintain service level objectives (SLOs) and error budgets to balance feature velocity with reliability.
  • Create tools and automation to eliminate manual toil and improve operational efficiency for engineering teams.
  • Advance Generative AI and Machine Learning capabilities within the reporting and technology PMO.
  • Manage infrastructure as code using tools like Terraform, CloudFormation, or Pulumi in AWS environments.
  • Develop production-quality code in Python, Go, or Node for automation and system integration.

What we're looking for

  • 7+ years of relevant experience.
  • Bachelor's degree in a relevant field of study or equivalent.
  • 5+ years of experience in site reliability engineering, systems engineering, or DevOps roles (preferred).
  • Deep expertise in AWS cloud and infrastructure as code tools like Terraform, CloudFormation, or Pulumi (preferred).
  • Experience defining/measuring SLIs, SLOs, and error budgets to drive reliability improvements (preferred).
  • Proficiency in AI development and strong programming skills in Python, Go, or Node (preferred).
  • Extensive experience with container orchestration (ECS or similar) and microservices architectures (preferred).
  • Proficiency with observability tools like Dynatrace, Prometheus, Grafana, Datadog, or New Relic (preferred).

More like this

Similar roles

Site Reliability Engineering Lead

US Bank

Atlanta, GA +2 11 days ago $111,605$131,300
SRE DevOps AWS Azure Kubernetes Docker Terraform Ansible Python PowerShell Shell Scripting CI/CD GitHub Actions Azure DevOps Jenkins GitLab Datadog Splunk Dynatrace Grafana Prometheus CloudWatch Azure Monitor OpenTelemetry ServiceNow Jira REST APIs SQL
6+ yrs exp

Senior Site Reliability Engineer

Salesforce

Remote (San Francisco, CA) 66 days ago $148,500$223,900
SRE Python Go Docker Kubernetes CI/CD Prometheus Grafana ELK Splunk Datadog Temporal Airflow Argo Workflows AWS GCP Linux Unix LLM Prompt Engineering
5+ yrs exp Remote

SW Engineer Developer Systems Reliability Engineering

Visa

Austin, TX 57 days ago $88,000$136,900
SRE CI-CD Prometheus Grafana Python Go Terraform Ansible Helm MySQL PostgreSQL MongoDB Splunk Datadog New Relic DynaTrace Sentry Azure DevOps Jenkins ArgoCD Cloud Formation Linux Windows
3+ yrs exp Hybrid

Site Reliability Engineer

Apple Inc

San Diego, CA 2 days ago $142,300$263,300
AWS Kubernetes Terraform Python Docker Argo Prometheus Elasticsearch Redis RDS ELB Linux Microservices Hybrid Cloud Networking
3+ yrs exp

Principal Site Reliability Engineer

Oracle

Nashville, TN 64 days ago $84,900$209,500
Site Reliability Engineering Oracle Cloud Infrastructure Linux Windows Server Python PowerShell Bash Ansible Chef infrastructure-as-code Networking DNS Firewalls Load Balancing Certificates incident-management Observability Capacity Planning
3+ yrs exp

Principal Site Reliability Engineer

Oracle

Reston, VA +1 52 days ago $84,900$209,500
Kubernetes Terraform Docker Python Bash Linux Unix Oracle Database RAC Chef Puppet DNS DHCP HTTP TCP/IP LLM VMware Cisco
6+ yrs exp