Site Reliability Engineer, Lead

Booz Allen Hamilton

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Chantilly, VA
Salary
$99,000–$225,000 / yr
Posted
49 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $180k
This role $162k
$84k most similar roles pay here $241k

This role pays less than 68% of similar roles. Most pay $145,375–$214,000 — the shaded band above. At the midpoint, this role pays about $162k versus about $180k for comparable roles.

Based on 238 similar postings.

Employer

About Booz Allen Hamilton

Booz Allen Hamilton is a management and technology consulting firm that provides analytics, digital, engineering, and cybersecurity solutions primarily to U.S. government agencies and commercial clients. Industry: Management & Technology Consulting

Booz Allen Hamilton currently has 802 open roles on FindRole.

Listed pay typically runs $86,800–$198,000 across 783 roles with salary data.

Most-posted roles

View all roles at Booz Allen Hamilton

At a glance

TL;DR · Site Reliability Engineer, Lead

As a Site Reliability Engineer, Lead, you will ensure the reliability, performance, scalability, and security of critical production systems and platforms across cloud and air-gapped environments. You will lead the design and implementation of observability, automation, incident response, and operational best practices while partnering with DevOps, infrastructure, and security teams to reduce risk. Your daily work involves performing root cause analysis, capacity planning, and developing technical tools to debug deployment issues within specific platforms. The role requires expertise in Linux systems administration, networking fundamentals within AWS, Python scripting, and Infrastructure as Code using Terraform and Terragrunt. You will utilize Prometheus, Grafana, the ELK stack, and Kubernetes for managing containerized workloads. Essential skills include experience with OpenTelemetry, Jenkins, Git, Docker, and knowledge of distributed systems to maintain highly available services in complex environments.

What you'll do

  • Ensure the reliability, performance, scalability, and security of critical production systems and platforms.
  • Design and implement observability, automation, incident response, and operational best practices across cloud and air-gapped environments.
  • Conduct root cause analysis and perform capacity planning to ensure high availability for services.
  • Develop technical tools to debug problems occurring during the deployment of applications within specific platforms.
  • Evaluate the health, stability, and reliability of systems in partnership with development and infrastructure teams.
  • Implement SRE practices including SLOs, SLIs, error budgets, and incident management protocols.
  • Manage Kubernetes environments and oversee containerized workloads across distributed microservices architectures.

What we're looking for

  • TS/SCI clearance with polygraph is required.
  • Must have 8+ years of experience with monitoring, logging, and observability platforms like Prometheus, Grafana, and ELK stack.
  • Must have 8+ years of experience with Linux systems administration and networking fundamentals within AWS.
  • Experience with Python scripting and automation is required.
  • Experience with Infrastructure as Code using Terraform and Terragrunt is required.
  • Knowledge of Kubernetes administration, troubleshooting, and operations is required.
  • A Bachelor's degree and 8+ years of experience in SRE, DevOps, or Platform Engineering are required, or 12+ years of experience in those fields without a degree.
  • Ability to obtain a Security+, SSCP, CCNA-Security, or GSEC certification within 6 months of start date.

More like this

Similar roles

Site Reliability Engineer

Balyasny Asset Management

Warsaw, Poland 78 days ago
Prometheus Grafana Loki Tempo OTEL Kubernetes Docker AWS Python Bash Go CI/CD DevOps SRE Agile
5+ yrs exp

Lead Site Reliability Engineer

Alloy

New York, NY 158 days ago $179,000$250,000
Kubernetes AWS Terraform Python Go Docker Infrastructure as Code Datadog CloudWatch ELK EFK Distributed Systems SRE
10+ yrs exp Hybrid

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 43 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 43 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 43 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp