Big Data/PySpark Engineering Lead Vice President

Citi

Confirmed live yesterday High trust

Quick summary

Work type
On-site
Location
Tampa, FL
Salary
$113,840–$170,760 / yr
Posted
15 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Below market

How this pay compares to similar roles

Similar $186k
This role $142k
$102k most similar roles pay here $227k

This role pays less than 89% of similar roles. Most pay $161,287–$211,200 — the shaded band above. At the midpoint, this role pays about $142k versus about $186k for comparable roles.

Based on 240 similar postings.

Employer

About Citi

Citi is one of the world’s most trusted financial institutions, proudly serving millions of customers across the United States.

Citi currently has 256 open roles on FindRole.

Listed pay typically runs $140,080–$210,120 across 238 roles with salary data.

Most-posted roles

View all roles at Citi

At a glance

TL;DR · Big Data/PySpark Engineering Lead Vice President

Big Data/PySpark Engineering Lead - Vice President is a senior-level role within the Applications Development team focused on leading systems analysis and programming activities. The successful candidate will design scalable, fault-tolerant batch and real-time data processing pipelines while migrating legacy data from on-premises SQL Servers to a modern Data Lakehouse environment. Key responsibilities include re-engineering complex ETL jobs into distributed frameworks, performing parity testing, and managing phased cutovers to ensure zero downtime. The role requires expert proficiency in Python, SQL, and Unix shell scripting, alongside experience with Spark, Hive, Kafka, Trino, Starburst, and NoSQL databases like MongoDB or HBase. Candidates must navigate data formats such as Parquet and Avro while utilizing tools like Bitbucket for source code management to solve complex problems involving legacy system decommissioning and modernizing infrastructure for high-volume financial data processing.

What you'll do

  • Design and implement scalable, fault-tolerant batch and real-time data processing pipelines using Spark, Flink, and Kafka.
  • Lead the migration of legacy data and logic from on-premises systems to a modern Data Lakehouse environment.
  • Re-engineer complex legacy ETL jobs into distributed processing frameworks using Python and Starburst/Trino.
  • Develop automated frameworks for data parity testing to ensure accuracy between legacy outputs and new big data results.
  • Write high-performance Python code and optimize complex SQL queries to reduce latency and costs.
  • Build and maintain CI/CD pipelines for the automated testing and deployment of data jobs.
  • Provide technical mentorship, conduct code reviews, and manage work allocation for junior and mid-level engineers.
  • Translate complex business requirements into detailed technical specifications for development teams.

What we're looking for

  • Bachelor's degree or equivalent experience is required, with a Master's degree preferred.
  • 6+ years of relevant experience in the Financial Service industry.
  • Experience as an Applications Development Manager or in a senior-level Applications Development role.
  • Expertise in Data Engineering focused on Big Data ecosystems including Hadoop, YARN, Hive, Impala, Spark, and Spark SQL.
  • Expert-level programming skills in Python and hands-on experience with Unix-based operating systems and shell scripting.
  • Proficiency with data formats (Avro, Parquet, CSV, JSON), NoSQL databases, and query engines like Trino, Presto, or Starburst.
  • Experience with source code management tools such as Bitbucket or Git.
  • Demonstrated leadership, project management, and stakeholder communication skills.

More like this

Similar roles

Big Data PySpark Lead Engineer Vice President

Citi

Remote (Jersey City, NJ) 42 days ago $142,320$213,480
PySpark Big Data Hadoop Hive HDFS Sqoop Spark Impala Scala SQL Shell Scripting Autosys Apache Kafka Distributed Systems
6+ yrs exp Remote

Big Data Support Engineer Assistant Vice President

Citi

Remote (Irving, TX) 128 days ago
Hadoop HDFS Hive Impala Spark YARN Sentry Oozie Kafka SQL Oracle Sybase Java C# .Net Shell Scripting Perl C++ HTML5 Linux Windows ITIL Problem Management Change Management Release Management uDeploy HP Diagnostics ITRS
5+ yrs exp Remote

Senior Data Engineer Assistant Vice President

Citi

Remote (Irving, TX) 115 days ago
Python Java Scala Spark Hadoop Kafka Hive Databricks SQL ETL ELT Microservices Docker Kubernetes Data Mesh Machine Learning Natural Language Processing Parquet Avro
8+ yrs exp Remote

Lead Data Engineer

Capital One Financial

McLean, VA +1 52 days ago $197,300$225,100
Python PySpark AWS SQL Scala Java Redshift Snowflake NoSQL Hadoop Hive EMR Kafka Spark MySQL Cassandra UNIX Linux Agile
4+ yrs exp

Lead Software Engineer - Data Engineer

JPMorgan Chase

Houston, TX 35 days ago
Python Java LLM RAG AI ML Kafka Redis Memcached Airflow Temporal Spark PySpark Databricks Snowflake Apache Iceberg AWS Azure GCP Dynatrace Splunk Grafana Microservices API Design
5+ yrs exp

Technology Lead Python Data Engineer Vice President

Citi

Remote (Rutherford, NJ) 14 days ago $142,320$213,480
Python SQL FastAPI Pydantic Django Docker Kubernetes Microservices CI/CD Linux Shell Scripting TDD Spark PySpark Hadoop Hive Neo4j LLMs GenAI Prompt Engineering
6+ yrs exp Remote