Big Data PySpark Lead Engineer Vice President

Citi

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Jersey City, NJ
Salary
$142,320–$213,480 / yr
Posted
42 days ago
Freshness
Confirmed live yesterday
Closes
Sep 30, 2026

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $196k
This role $178k
$132k most similar roles pay here $235k

This role pays less than 67% of similar roles. Most pay $174,462–$216,656 — the shaded band above. At the midpoint, this role pays about $178k versus about $196k for comparable roles.

Based on 240 similar postings.

Employer

About Citi

Citi is one of the world’s most trusted financial institutions, proudly serving millions of customers across the United States.

Citi currently has 256 open roles on FindRole.

Listed pay typically runs $140,080–$210,120 across 238 roles with salary data.

Most-posted roles

View all roles at Citi

At a glance

TL;DR · Big Data PySpark Lead Engineer Vice President

The Big Data PySpark Lead Engineer - Vice President is a senior-level role within the Technology team focused on leading applications systems analysis and programming activities. The successful candidate will build and maintain scalable data pipelines using PySpark to process large volumes of structured and unstructured data while designing solutions across the Hadoop ecosystem, including Hive, HDFS, Sqoop, Spark, Impala, and Scala. Day-to-day responsibilities involve managing real-time and batch workflows, writing complex SQL queries for distributed systems, and automating pipeline scheduling using shell scripting and Autosys. The role requires expertise in data modeling, warehouse principles, and dimensional modeling to ensure consistency across the platform. This position addresses critical technical challenges in data ingestion and processing within a regulated financial services environment where high availability, low-latency delivery, and rigorous data governance are essential for maintaining system integrity.

What you'll do

  • Build and maintain scalable data pipelines using PySpark to process large volumes of structured and unstructured data.
  • Develop solutions across the Hadoop ecosystem, including Hive, HDFS, Sqoop, Spark, Impala, and Scala.
  • Manage real-time and batch data workflows using streaming platforms to ensure high availability and low latency.
  • Write complex SQL queries to extract, validate, and analyze data across distributed systems.
  • Design and implement data models and architecture patterns aligned with data warehouse principles.
  • Automate pipeline scheduling and orchestration using shell scripting and Autosys tools.
  • Identify, assess, and resolve technical risks and data issues to maintain system integrity.
  • Serve as a coach and advisor to mid-level developers while allocating work as necessary.

What we're looking for

  • Bachelor's degree or equivalent experience is required, with a Master's degree preferred.
  • 6 to 10 years of relevant experience in applications development or systems analysis roles.
  • Hands-on expertise in PySpark and building distributed data workflows at scale.
  • Practical knowledge of the Hadoop ecosystem including Hive, HDFS, Sqoop, Spark, Impala, and Scala.
  • Proficiency in complex SQL query development for data analysis and transformation across large datasets.
  • Competence in shell scripting and job scheduling using Autosys or equivalent automation tools.
  • Demonstrated leadership, project management skills, and the ability to coach mid-level developers.
  • Strong communication skills to articulate technical concepts to both technical and non-technical audiences.

More like this

Similar roles

Lead Data Engineer

Capital One Financial

McLean, VA +1 52 days ago $197,300$225,100
Python PySpark AWS SQL Scala Java Redshift Snowflake NoSQL Hadoop Hive EMR Kafka Spark MySQL Cassandra UNIX Linux Agile
4+ yrs exp

Big Data Support Engineer Assistant Vice President

Citi

Remote (Irving, TX) 128 days ago
Hadoop HDFS Hive Impala Spark YARN Sentry Oozie Kafka SQL Oracle Sybase Java C# .Net Shell Scripting Perl C++ HTML5 Linux Windows ITIL Problem Management Change Management Release Management uDeploy HP Diagnostics ITRS
5+ yrs exp Remote

Lead Data Engineer

Capital One Financial

Richmond, VA 3 days ago $179,400$204,700
Python Spark PySpark AWS Redshift Snowflake Java Scala NoSQL ETL Kafka Hadoop Hive EMR DynamoDB Cassandra Agile
4+ yrs exp

Data Analytics Lead Engineer

Citi

Irving, TX 9 days ago $125,760$188,640
SQL Python Scala Apache Spark PySpark Databricks Snowflake Hadoop Delta Lake Hive Impala Iceberg Apache Airflow Autosys Tableau Cognos NoSQL MongoDB Cassandra AWS Azure GCP
6+ yrs exp Hybrid

Senior Data Engineer Vice President

Citi

Remote (Irving, TX) 51 days ago $125,760$188,640
Hadoop Apache Kafka Spark PySpark Python SQL FastAPI MLOps AI/ML GenAI AWS Azure Google Cloud ETL ELT Hive HDFS YARN MapReduce HBase Unix
7+ yrs exp Remote