This listing doesn't post a salary. Most similar roles pay
$151,587–$215,881.
Based on 239 similar postings.
Employer
About JPMorgan Chase
JPMorgan Chase & Co. is a global financial services firm and one of the largest banks in the world, offering investment banking, commercial banking, asset management, and consumer financial services.
JPMorgan Chase currently has
1137 open roles
on FindRole.
As a Technology Support Lead - Problem Management & Governance Lead within the Production Management team, you will serve as a subject matter expert for end-to-end problem management across application and infrastructure services. You will identify root causes, analyze incident trends, and manage remediation actions to improve service stability and reduce recurring issues. Your daily work involves facilitating structured root cause analyses, maintaining known-error records, and ensuring all actions meet risk, control, and resiliency standards. You will leverage enterprise-authorized AI capabilities to enhance triage speed while ensuring human-in-the-loop validation for sensitive data. Key requirements include expertise in ITIL frameworks, observability tools, and managing both on-premises and public cloud environments. You must possess strong analytical skills to translate operational data into actionable improvements while influencing stakeholders through clear communication regarding technical findings and risk mitigation.
Manage the end-to-end problem lifecycle including identification, root cause analysis, remediation, and closure for application and infrastructure services.
Analyze incident trends and operational data to identify systemic issues and prioritize remediation based on business impact and risk.
Facilitate structured root cause analyses and post-incident reviews to document lessons learned and implement corrective actions.
Integrate enterprise-authorized AI capabilities into workflows to improve incident triage speed and consistency while ensuring data security.
Establish ownership, milestones, and success measures for problem records while providing transparent reporting to stakeholders.
Maintain accurate problem records, known-error databases, and technical findings aligned with established standards.
Identify operational risks and control deficiencies to coordinate mitigation, governance, and compliance actions.
Define and report key performance metrics regarding remediation effectiveness, incident volume reduction, and root cause categories.
What we're looking for
5+ years of experience or equivalent expertise troubleshooting, resolving, and maintaining information technology services.
Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support production operations workflows with strong validation habits.
Experience managing applications or infrastructure in a large-scale technology environment both on premises and public cloud.
Proficiency in observability and monitoring tools and techniques.
Experience executing processes within the Information Technology Infrastructure Library (ITIL) framework.
Demonstrated experience managing end-to-end problem records, root cause analyses, corrective actions, known errors, and recurrence-prevention activities.
Strong analytical skills to identify systemic problems and prioritize remediation based on impact and risk.
Proven stakeholder management and influencing skills to communicate effectively with senior business and technology stakeholders.
Experience supporting Problem Management, Production Management, Site Reliability Engineering, or Operational Excellence in a large enterprise (preferred).
Knowledge of structured root cause analysis methods, post-incident review practices, and corrective/preventive action management (preferred).
Experience defining and reporting problem management metrics such as recurrence, aging, and remediation effectiveness (preferred).
Familiarity with enterprise-authorized AI, automation, analytics, or AIOps capabilities for operational analysis (preferred).
Relevant industry certification or formal training in ITIL, Service Management, SRE, cloud technology, or risk management (preferred).
Site Reliability Engineering
SRE
Incident Management
Problem Management
Observability
Automation
AI
IT Service Management
Agile
Lean
Service Level Indicators
Service Level Objectives
Error Budgets
Program Management