Vice President, Engineering SRE Platforms Site Reliability Engineer

Goldman Sachs

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Dallas, TX
Posted
75 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

How this pay compares to similar roles

Similar $184k
$135k most similar roles pay here $235k

This listing doesn't post a salary. Most similar roles pay $154,323–$214,125.

Based on 240 similar postings.

Employer

About Goldman Sachs

Goldman Sachs is a leading global investment banking, securities, and investment management firm providing financial services to corporations, financial institutions, governments, and individuals.

Goldman Sachs currently has 134 open roles on FindRole.

Listed pay typically runs $137,000–$250,000 across 55 roles with salary data.

Most-posted roles

View all roles at Goldman Sachs

At a glance

TL;DR · Vice President, Engineering SRE Platforms Site Reliability Engineer

Engineering - SRE Platforms - Site Reliability Engineer - Vice President joins the Engineering Division to ensure the availability, reliability, and scalability of critical platform services across on-premises datacenters and multiple public cloud environments. This leader will architect and build fault-tolerant systems, manage complex incident responses, perform root cause analyses, and oversee capacity planning while mentoring senior engineers. The role involves developing automated tools for deployment, monitoring, and logging to eliminate toil and improve operational workflows. Key technical requirements include proficiency in Java, Python, or Go; experience with AWS, GCP, Docker, and Kubernetes; mastery of Infrastructure as Code via Terraform or CloudFormation; and expertise in Prometheus, Grafana, and the ELK stack. Additionally, the role requires advanced skills in Prompt Engineering and RAG architectures to automate SRE workflows within a distributed systems environment.

What you'll do

  • Drive the strategic direction for availability, scalability, and performance of mission-critical platform services.
  • Lead the design and implementation of highly available, resilient, and scalable infrastructure architectures.
  • Develop sophisticated automation tools and platforms to eliminate toil and optimize operational workflows.
  • Lead critical incident responses and conduct in-depth root cause analysis for systemic issues.
  • Partner with development teams to integrate reliability into application designs and manage capacity planning.
  • Implement advanced monitoring, logging, and tracing strategies to provide actionable insights into system health.
  • Provide technical vision, perform code reviews, and mentor senior and staff-level engineers.
  • Evaluate and integrate cutting-edge technologies to improve operational efficiency and firm-wide reliability.

What we're looking for

  • Minimum of 6 years of hands-on experience in Site Reliability Engineering for enterprise-level systems.
  • Advanced degree in Computer Science or a related technical field, or equivalent practical experience.
  • Exceptional programming skills in major languages such as Java, Python, or Go.
  • Extensive experience with cloud platforms (AWS, GCP), containerization (Docker, Kubernetes), and Infrastructure as Code tools.
  • Advanced proficiency in Prompt Engineering and Retrieval-Augmented Generation (RAG) architectures to automate SRE workflows.
  • Expertise in Linux internals, networking, distributed systems, and performance tuning.
  • Proficiency in monitoring, alerting, logging, tracing, and CI/CD tools.
  • Experience with Distributed Databases like Elastic Search, GCP Big Query, or messaging systems like Kafka (preferred).

More like this

Similar roles

Vice President, Site Reliability Engineering

Goldman Sachs

Dallas, TX 75 days ago
Java Python Node.js Terraform Ansible CloudFormation Docker Kubernetes AWS GCP Azure Prometheus Grafana Splunk Datadog OpenTelemetry ELK CloudWatch Linux
7+ yrs exp

SRE Platforms Engineer Associate

Goldman Sachs

Dallas, TX 18 days ago
Site Reliability Engineering Go Python C C++ Java Perl Ruby Shell Scripting Unix Networking Distributed Systems Cloud Platforms AI Tools Automation Data Structures Algorithms System Design
2+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Palo Alto, CA 59 days ago
Site Reliability Engineering Java Go Python Terraform Kubernetes Docker CI/CD GitOps Grafana Prometheus Dynatrace Datadog Splunk Kafka RabbitMQ SQS Neo4j Pinecone Weaviate Chroma LangChain LangGraph AutoGen CrewAI GitHub Copilot Fluentd Logstash Vector RESTful APIs RAG TensorFlow PyTorch scikit-learn Hadoop Spark Flink MongoDB Cassandra DynamoDB InfluxDB TimescaleDB AWS Azure GCP
5+ yrs exp

Senior Site Reliability Engineer

Fiserv

Berkeley Heights, NJ 6 days ago $128,000$216,000
AWS Kubernetes Terraform CI/CD GitHub Actions Python Bash Ruby on Rails Docker Linux Unix RDBMS Document Storage New Relic Dynatrace Datadog DNS Load Balancing Virtual Networking