Principal Reliability Engineer

The Hartford

Confirmed live yesterday High trust
Hybrid

Quick summary

Work type
Hybrid
Location
Columbus, OHChicago, ILHartford, CTCharlotte, NC
Salary
$152,800–$229,200 / yr
Posted
43 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $190k
This role $191k
$133k most similar roles pay here $245k

This role pays more than 52% of similar roles. Most pay $165,700–$214,625 — the shaded band above. At the midpoint, this role pays about $191k versus about $190k for comparable roles.

Based on 240 similar postings.

Employer

About The Hartford

The Hartford is a leading provider of property and casualty insurance, group benefits, and mutual funds, serving businesses and individuals across the United States. Industry: Insurance & Financial Services

The Hartford currently has 45 open roles on FindRole.

Listed pay typically runs $127,600–$191,400 across 37 roles with salary data.

Most-posted roles

View all roles at The Hartford

At a glance

TL;DR · Principal Reliability Engineer

Principal Reliability Engineer - EDS serves as the senior technical authority within the Enterprise Data Services organization, overseeing the reliability, resilience, and performance of data platforms, cloud infrastructure, and pipelines. This role establishes the strategic vision for reliability engineering while leading the development of observability frameworks, automation-first operations, and AIOps initiatives. You will build automated remediation tools using LLMs and cloud-native AI services like AWS Bedrock and Vertex AI to manage complex systems. Key responsibilities include architecting fail-safe patterns for Snowflake, EMR, Hadoop/Spark, and Kubernetes environments while enforcing IaC and CI/CD standards. The role requires expertise in Python, Terraform, and CloudFormation to ensure data quality and pipeline reliability. You will solve critical challenges regarding large-scale distributed systems, ensuring high availability across multi-cloud environments including AWS and GCP for mission-critical data products.

What you'll do

  • Define and execute the reliability engineering strategy for data platforms, cloud environments, and data products.
  • Architect and manage high-availability infrastructure across AWS, GCP, Snowflake, and Hadoop/Spark clusters.
  • Develop AI-driven automation for anomaly detection, alert correlation, and predictive capacity management using LLMs and machine learning.
  • Design and implement enterprise-wide observability frameworks including logging, metrics, tracing, and distributed profiling.
  • Establish SLO/SLI frameworks and incident response patterns to ensure data quality and pipeline reliability.
  • Create automated tools for self-healing systems and the elimination of operational toil.
  • Set standards for Infrastructure as Code (IaC), CI/CD pipelines, and technical documentation across the organization.
  • Provide technical leadership and mentorship to senior engineers while influencing executive-level architectural decisions.

What we're looking for

  • Must have 10+ years of experience in data, cloud, platform engineering, site reliability engineering, or large-scale distributed systems.
  • Experience in leadership or technology leader roles is required.
  • Proficiency with data and cloud platforms including architectural patterns for resilience, networking, security, and distributed infrastructure.
  • Deep experience supporting or engineering Snowflake, EMR, Hadoop/Spark, and Data Integration platforms.
  • Proficiency in scripting and programming, preferably using Python, for automation and reliability frameworks.
  • Experience with Infrastructure-as-Code (Terraform, CloudFormation) and enterprise CI/CD systems.

More like this

Similar roles

Principal Site Reliability Engineer

Nvidia

Santa Clara, CA 3 days ago $248,000$396,750
Kubernetes Distributed Systems Python Go Terraform AWS Azure GCP OpenTelemetry infrastructure-as-code Linux TypeScript JavaScript Java Crossplane AWS CDK CloudFormation AI/ML Platforms High-Performance Computing
10+ yrs exp Hybrid

Lead Principal Site Reliability Engineer

Oracle

Nashville, TN 57 days ago $96,300$264,100
Site Reliability Engineering Kubernetes Docker Terraform Ansible Chef Puppet Python Go Java JavaScript Bash Oracle Cloud Infrastructure Microsoft Azure Google Cloud Platform infrastructure-as-code Chaos Engineering
6+ yrs exp

Lead Site Reliability Engineer

JPMorgan Chase

Plano, TX 44 days ago
Site Reliability Engineering Python Java Spring Boot .NET CI/CD Docker Kubernetes Terraform AWS Prometheus Grafana Dynatrace Datadog Splunk infrastructure-as-code CloudFormation ECS Incident Management Observability
5+ yrs exp

Senior Lead Site Reliability Engineer

JPMorgan Chase

Palo Alto, CA 59 days ago
Site Reliability Engineering Java Go Python Terraform Kubernetes Docker CI/CD GitOps Grafana Prometheus Dynatrace Datadog Splunk Kafka RabbitMQ SQS Neo4j Pinecone Weaviate Chroma LangChain LangGraph AutoGen CrewAI GitHub Copilot Fluentd Logstash Vector RESTful APIs RAG TensorFlow PyTorch scikit-learn Hadoop Spark Flink MongoDB Cassandra DynamoDB InfluxDB TimescaleDB AWS Azure GCP
5+ yrs exp

Senior Staff Principal Reliability Engineer

Qualcomm

San Diego, CA 67 days ago $180,200$270,200
Semiconductor Reliability BTI TDDB HCI Electromigration TSV Chiplets SoC Weibull Analysis JMP Minitab Statistical Analysis Lifetime Modeling FinFET GAA CMOS
8+ yrs exp