Vice President Site Reliability Engineering

Goldman Sachs

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
New York, NY
Posted
1 day ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $189k
$135k most similar roles pay here $245k

This listing doesn't post a salary. Most similar roles pay $161,437–$217,500.

Based on 238 similar postings.

Employer

About Goldman Sachs

Goldman Sachs is a leading global investment banking, securities, and investment management firm providing financial services to corporations, financial institutions, governments, and individuals.

Goldman Sachs currently has 96 open roles on FindRole.

Listed pay typically runs $130,000–$250,000 across 41 roles with salary data.

Most-posted roles

View all roles at Goldman Sachs

At a glance

TL;DR · Vice President Site Reliability Engineering

As Vice President - Site Reliability Engineering (SRE) – The Core Engineering, you will join a team focused on engineering highly reliable, observable, and resilient platforms for critical business services. You will collaborate with multiple engineering teams to improve production system architecture, facilitate fast service delivery, and reduce downtime by implementing SLOs, error budgets, and blameless post-mortems. Your daily work involves building automation tools to reduce operational toil, conducting architectural reviews for fault-tolerant systems, and leading responses to complex multi-system incidents. You will utilize Java, Python, or Node.js alongside Infrastructure as Code tools like Terraform, Ansible, or CloudFormation. The role requires expertise in Docker, Kubernetes, and major cloud providers like AWS, GCP, or Azure. You will also leverage observability stacks including Prometheus, Grafana, and Datadog to manage distributed systems within the financial markets domain.

What you'll do

  • Establish service level objectives (SLOs), service level indicators (SLIs), and error budgets with engineering leadership.
  • Architect highly available, fault-tolerant, and self-healing systems using patterns like circuit breakers and rate limiting.
  • Build automation, tooling, and self-service capabilities to reduce operational toil and manual work.
  • Improve production readiness through load testing, performance tuning, and capacity forecasting.
  • Lead responses to complex, multi-system production incidents and facilitate blameless post-mortems.
  • Design sustainable on-call models with clear escalation paths and balanced pager responsibilities.
  • Develop clean, maintainable code for infrastructure automation using IaC frameworks like Terraform or Ansible.
  • Manage observability stacks including distributed tracing, logging, and metrics to ensure system visibility.

What we're looking for

  • Proficiency in at least one major programming language such as Java, Python, or Node.js for tooling and automation.
  • Hands-on experience with Infrastructure as Code (IaC) frameworks like Terraform, Ansible, or CloudFormation.
  • Deep understanding of containerization and orchestration technologies including Docker, Kubernetes, service meshes, and ingress controllers.
  • Advanced experience building and operating highly resilient cloud-native architectures on major providers like AWS, GCP, or Azure.
  • Proficiency with observability stacks including distributed tracing, logging, and metrics using tools like Prometheus, Grafana, or Datadog.
  • Knowledge of networking protocols, load balancing strategies, and Linux environment development.
  • Ability to translate complex technical issues into actionable insights for both technical and non-technical stakeholders.
  • Bachelor’s degree in Computer Science, System Engineering, or a related technical field (preferred); 7 to 10 years of experience (preferred).

More like this

Similar roles

Vice President, Site Reliability Engineering

Goldman Sachs

Dallas, TX 82 days ago
Java Python Node.js Terraform Ansible CloudFormation Docker Kubernetes AWS GCP Azure Prometheus Grafana Splunk Datadog OpenTelemetry ELK CloudWatch Linux
7+ yrs exp

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 11 days ago $248,000$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK CloudFormation AI/ML Platforms High-Performance Computing
10+ yrs exp Hybrid

Site Reliability Engineering Lead

US Bank

Atlanta, GA +2 11 days ago $111,605$131,300
SRE DevOps AWS Azure Kubernetes Docker Terraform Ansible Python PowerShell Shell Scripting CI/CD GitHub Actions Azure DevOps Jenkins GitLab Datadog Splunk Dynatrace Grafana Prometheus CloudWatch Azure Monitor OpenTelemetry ServiceNow Jira REST APIs SQL
6+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Chicago, IL 16 days ago
SRE Python Ansible Terraform Kubernetes Docker AWS Prometheus Grafana Dynatrace Datadog Splunk Linux Windows Jenkins GitLab Network as Code
3+ yrs exp

Site Reliability Engineer III

JPMorgan Chase

Chicago, IL 16 days ago
SRE Python Ansible Terraform Kubernetes Docker AWS Prometheus Grafana Dynatrace Datadog Splunk Linux Windows Jenkins GitLab Network as Code
3+ yrs exp