Vice President, Site Reliability Engineering

Goldman Sachs

Confirmed live 2 days ago High trust

Quick summary

Work type
On-site
Location
Dallas, TX
Posted
75 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

How this pay compares to similar roles

Similar $180k
$117k most similar roles pay here $247k

This listing doesn't post a salary. Most similar roles pay $152,150–$207,625.

Based on 238 similar postings.

Employer

About Goldman Sachs

Goldman Sachs is a leading global investment banking, securities, and investment management firm providing financial services to corporations, financial institutions, governments, and individuals.

Goldman Sachs currently has 134 open roles on FindRole.

Listed pay typically runs $137,000–$250,000 across 55 roles with salary data.

Most-posted roles

View all roles at Goldman Sachs

At a glance

TL;DR · Vice President, Site Reliability Engineering

As Site Reliability Engineering (SRE), The Core Engineering, Vice President, you will join a team focused on engineering highly reliable, observable, and resilient platforms for critical business services. You will collaborate with multiple engineering teams to improve production system architecture, facilitate fast service delivery, and reduce downtime by implementing SLOs, error budgets, and blameless post-mortems. Your daily work involves building automation tools to reduce operational toil, conducting architectural reviews for fault-tolerant systems, and leading responses to complex multi-system incidents. You will utilize Java, Python, or Node.js to build tooling while leveraging Infrastructure as Code frameworks like Terraform, Ansible, or CloudFormation. The role requires expertise in Docker, Kubernetes, and major cloud providers like AWS, GCP, or Azure, alongside observability stacks including Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch within a distributed systems environment.

What you'll do

  • Establish service level objectives (SLOs), indicators (SLIs), and error budgets with engineering leadership.
  • Architect highly available, fault-tolerant, and self-healing systems using patterns like circuit breakers and rate limiting.
  • Build automation, tooling, and self-service capabilities to reduce manual operational toil.
  • Improve production readiness through load testing, performance tuning, and capacity forecasting.
  • Lead responses to complex multi-system incidents and facilitate blameless post-mortems to identify root causes.
  • Design sustainable on-call models with clear escalation paths and balanced pager responsibilities.
  • Develop clean, maintainable code for infrastructure automation using IaC frameworks like Terraform or Ansible.
  • Manage observability stacks including distributed tracing, logging, and metrics across cloud-native architectures.

What we're looking for

  • Bachelor's degree in Computer Science, System Engineering, or a related technical field involving programming.
  • 7 to 10 years of experience.
  • Proficiency in at least one major programming language such as Java, Python, or Node.js.
  • Hands-on experience with Infrastructure as Code frameworks like Terraform, Ansible, or CloudFormation.
  • Deep understanding of containerization and orchestration technologies including Docker and Kubernetes.
  • Advanced experience building and operating resilient cloud-native architectures on AWS, GCP, or Azure.
  • Proficiency with observability stacks including distributed tracing, logging, and metrics tools.
  • Knowledge of networking protocols, load balancing, and Linux environment development.

More like this

Similar roles

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 3 days ago $248,000$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK CloudFormation AI/ML Platforms High-Performance Computing
10+ yrs exp Hybrid

Site Reliability Engineering Lead

US Bank

Atlanta, GA +2 3 days ago $111,605$131,300
SRE DevOps AWS Azure Kubernetes Docker Terraform Ansible Python PowerShell Shell Scripting CI/CD GitHub Actions Azure DevOps Jenkins GitLab Datadog Splunk Dynatrace Grafana Prometheus CloudWatch Azure Monitor OpenTelemetry ServiceNow Jira REST APIs SQL
6+ yrs exp

Senior Manager Site Reliability Engineer

The Walt Disney Company

Remote 16 days ago $175,000$215,000
SRE AWS GCP Azure Kubernetes Terraform Ansible Harness GitLab CloudFormation CI/CD Observability Automation Serverless DevOps
10+ yrs exp Remote

Senior Site Reliability Engineer

CVS Health

Woonsocket, RI +1 14 days ago $92,700$203,940
SRE DevOps Kubernetes OpenShift Docker CI/CD Splunk Dynatrace Datadog Prometheus Grafana Python Java AWS Microsoft Azure Google Cloud Rancher GitHub BitBucket Jenkins Microservices web API’s Apigee
5+ yrs exp Hybrid