Principal System Engineering, SRE

AT&T

Confirmed live yesterday High trust
Closes in 7 days

Quick summary

Work type
On-site
Location
Atlanta, GA
Salary
$155,400–$261,100 / yr
Posted
8 days ago
Freshness
Confirmed live yesterday
Closes
Sep 18, 2026 (soon)

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $185k
This role $208k
$129k most similar roles pay here $275k

This role pays more than 57% of similar roles. Most pay $156,600–$214,125 — the shaded band above. At the midpoint, this role pays about $208k versus about $185k for comparable roles.

Based on 240 similar postings.

Employer

About AT&T

AT&T is a US-based telecommunications company providing wireless, broadband, and fiber internet service along with phone and connectivity products for consumers and businesses.

AT&T currently has 61 open roles on FindRole.

Listed pay typically runs $117,350–$226,800 across 56 roles with salary data.

Most-posted roles

View all roles at AT&T

At a glance

TL;DR · Principal System Engineering, SRE

As a Principal System Engineering - SRE within the Systems Reliability and Software Delivery team, you will focus on identifying why production incidents occur and implementing long-term preventive measures. You will perform deep root cause analysis across applications, infrastructure, and cloud environments by analyzing observability data to uncover systemic weaknesses and patterns. Your daily work involves creating structured postmortems and partnering with engineering teams to drive corrective actions. To succeed, you must possess expertise in distributed systems, end-to-end architecture, and tools such as T-APM, T-Trace, CatchPoint, Grafana, Python, SQL, Power BI, and ServiceNow. You will utilize AI-assisted analysis and data visualization to transition the organization from reactive responses to proactive reliability. The role requires proficiency in Agile, DevOps, CI/CD, and modern Release Management within a complex enterprise environment.

What you'll do

  • Analyze production incidents end-to-end across applications, infrastructure, and cloud environments.
  • Use observability data to identify root causes, patterns, and systemic weaknesses.
  • Create high-quality postmortems based on incident insights.
  • Partner with engineering teams to implement permanent fixes and preventive improvements.
  • Utilize automation and AI-assisted analysis to transition from reactive response to proactive reliability.
  • Analyze operational data using tools, queries, or advanced analytics for pattern detection.
  • Drive long-term improvements by identifying recurring issues in the system architecture.

What we're looking for

  • Must have 7+ years of experience in Systems Engineering, ITSM, or Release/Change Management.
  • Experience in SRE, Support, or QA roles is required.
  • Proficiency with observability tools such as T-APM, T-Trace, CatchPoint, or Grafana is required.
  • Experience with Python, SQL, Power BI, and ITSM tools like ServiceNow is required.
  • Must have experience with SAFe, Agile, DevOps, CI/CD, and building Gen AI use cases.
  • Ability to perform deep root cause analysis for production incidents across cloud, web, and infrastructure environments.
  • Experience using AI or advanced analytics for incident analysis and pattern detection.
  • A Bachelor’s degree in Computer Science and relevant certifications (SAFe, Agile, DevOps, AI/ML) are preferred.

More like this

Similar roles

Principal System Engineering, SRE

AT&T

Plano, TX 8 days ago $155,400$261,100
SRE Python SQL Power BI Tableau Grafana CatchPoint ServiceNow Jira Cloud Git CI/CD Agile SAFe DevOps Data Analytics Gen AI ITSM Release Management Change Management Distributed Systems
7+ yrs exp

Principal System Engineering, SRE

AT&T

Dallas, TX +1 8 days ago $155,400$261,100
SRE Python SQL Power BI Tableau Grafana CatchPoint ServiceNow Jira Git CI/CD Agile SAFe DevOps Gen AI Data Analytics ITSM Release Management Change Management Distributed Systems
7+ yrs exp

Principal System Engineering

AT&T

Alpharetta, GA +4 9 days ago $155,400$261,100
Kafka IXBUS Microservices CI/CD Observability Monitoring Logging Cybersecurity Disaster Recovery Capacity Planning Cost Optimization Automation Event-Driven Architecture
7+ yrs exp

Principal System Engineering

AT&T

Atlanta, GA +5 9 days ago $155,400$261,100
Kafka IXBUS Microservices CI/CD Observability Disaster Recovery Capacity Planning Cost Optimization Automation
7+ yrs exp

Principal System Engineering

AT&T

Alpharetta, GA +4 9 days ago $155,400$261,100
Kafka IXBUS Microservices CI/CD Observability Monitoring Logging Cybersecurity Disaster Recovery Capacity Planning Cost Optimization Automation Event-Driven Architecture
7+ yrs exp