Lead Site Reliability Engineer

Alloy

Confirmed live 2 days ago High trust
Hybrid

Quick summary

Work type
Hybrid
Location
New York, NY
Salary
$179,000–$250,000 / yr
Posted
158 days ago
Freshness
Confirmed live 2 days ago

Market check

Salary context

Above market

How this pay compares to similar roles

Similar $180k
This role $214k
$129k most similar roles pay here $263k

This role pays more than 78% of similar roles. Most pay $147,200–$213,125 — the shaded band above. At the midpoint, this role pays about $214k versus about $180k for comparable roles.

Based on 238 similar postings.

Employer

About Alloy

Alloy is an identity decisioning platform that provides fraud prevention, compliance, and credit underwriting solutions for banks, credit unions, and fintechs to automate identity verification decisions. Industry: Financial Technology & Identity Verification

Alloy currently has 23 open roles on FindRole.

Listed pay typically runs $155,000–$190,000 across 23 roles with salary data.

Most-posted roles

View all roles at Alloy

At a glance

TL;DR · Lead Site Reliability Engineer

As a Lead Site Reliability Engineer on the Infrastructure team, you will manage a large footprint of Kubernetes clusters and databases while transforming complex, fragile systems into automated, self-service platforms. You will design and build systems to automate infrastructure management, including provisioning, upgrades, and migrations, while reducing operational toil by creating reliable workflows. Your daily work involves building internal tooling for other engineers, improving the reliability of distributed systems, and contributing to architecture decisions regarding infrastructure, reliability, and security. To succeed, you must possess strong software engineering skills and experience with Infrastructure as Code using Terraform. You will utilize tools such as Docker, Kubernetes, Datadog, CloudWatch, and ELK/EFK while writing production-quality code in languages like Python, Go, or JavaScript to solve problems related to scaling and securing large-scale data systems.

What you'll do

  • Design and build systems to automate infrastructure management including provisioning, upgrades, and migrations.
  • Convert manual operational processes into reliable, repeatable automated workflows to reduce toil.
  • Build internal tools and platforms that enable safe self-service changes for other engineers.
  • Improve the reliability and resilience of Kubernetes clusters, databases, and distributed services.
  • Implement and evolve systems for deploying and running applications in Kubernetes environments.
  • Contribute to architectural decisions regarding infrastructure, reliability, and security.
  • Write and review production-quality code to build robust systems rather than simple scripts.
  • Participate in on-call rotations while building proactive systems to prevent incidents from occurring.

What we're looking for

  • 10+ years of experience in infrastructure, SRE, or software engineering roles.
  • Strong software engineering skills to build systems rather than just scripts.
  • Experience managing production infrastructure at scale using cloud and containerized systems.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Experience running and troubleshooting distributed systems using Docker and Kubernetes.
  • Proficiency in at least one programming language including Python, Go, or JavaScript.
  • Experience with observability and debugging tools like Datadog, CloudWatch, or ELK/EFK.
  • Strong communication and collaboration skills to work within a small team.

More like this

Similar roles

Staff Site Reliability Engineer, Kubernetes

Okta Inc

Bellevue, WA +4 15 days ago $194,000$267,000
Kubernetes AWS Helm Karpenter Istio Terraform CI/CD Python Bash Go Docker Prometheus Grafana CloudWatch ELK Stack S3 RDS EC2 IAM
Hybrid

Lead Site Reliability Engineer

JPMorgan Chase

New York, NY 43 days ago
Site Reliability Engineering SRE CI/CD Observability Monitoring Automation Change Management Dynatrace Splunk Geneos Grafana ITIL AWS Azure GCP Python Shell PowerShell Ansible Terraform Kubernetes OpenShift Microservices
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 43 days ago
SRE CI/CD Jenkins GitLab Terraform Docker Kubernetes ECS AI Python Go JavaScript GraphQL Kafka OpenTelemetry Networking
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 43 days ago
Site Reliability Engineering Python Java Spring Boot .Net Kubernetes AWS Google Cloud CI/CD Observability Monitoring Telemetry Networking Infrastructure Optimization FinOps Disaster Recovery Capacity Management
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Jersey City, NJ 45 days ago
Site Reliability Engineering Python Java Spring Boot .Net AI CI/CD Container Orchestration AWS Observability Monitoring Telemetry Networking System Architecture SDLC
5+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp