Citi

Citi

Big Data/PySpark Engineering Lead - Vice President

Tampa Florida United States · Exec · Full-time

Sponsorship not specified$114k-$171kDetected 2 days ago
PythonData StructuresAlgorithmsCode ReviewGitSQLNoSQLMongoDBCI/CDLinuxDevOpsKafkaSparkData EngineeringTest AutomationLeadershipMentoringHadoopBitbucket

About the role

  • ETL/ELT Transformation: Re-engineer existing stored procedures and complex legacy ETL jobs into scalable, distributed processing frameworks using Spark (Python) and Starburst/Trino.
  • Schema Evolution: Map and transform rigid, legacy relational schemas into flexible, high-performance formats optimized for the cloud (e.g., Parquet, Avro, or Iceberg).

Responsibilities

  • Architecture & Design Design and implement scalable, fault-tolerant batch and real-time data processing pipelines.
  • Develop robust data models and schema designs optimized for both performance and storage efficiency.
  • Validation & Parity Testing: Design and implement automated frameworks for Data Parity Testing to ensure 100% accuracy and consistency between legacy outputs and new big data results.
  • Optimize complex SQL queries and fine-tune distributed computing clusters to reduce latency and costs.
  • DevOps & Reliability Build and maintain CI/CD pipelines for automated testing and deployment of data jobs.
  • Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements

Requirements

  • Knowledge in Hadoop, YARN, Hive, Impala, Spark, and Spark SQL with extensive high volume of data processing pipeline development.
  • Programming Expert level and hand on experience in Python.
  • Familiarity with data formats like Avro, Parquet, CSV, JSON.
  • Hands-on experience in writing SQL queries.
  • Experience with source code management tools such as Bitbucket, Git etc.
  • Big Data Tech Proficiency and hands-on in Hadoop, Spark, Hive, Kafka, and NoSQL databases (MongoDB, HBase).
  • Experience working with query engines like Trino, Presto, Starburst Strong computer science fundamentals in data structures, algorithms, databases, and operating systems.
  • 6+ years of relevant experience in the Financial Service industry Experience as Applications Development Manager Experience as senior level in an Applications Development role Experience in Data Engineering, focused on Big Data ecosystems.
  • Highly experienced with Unix based operating systems and shell scripting.
  • Reverse Engineering, ability to read "spaghetti" SQL or old scripts and document the business logic before moving it.
  • Data Lineage, Experience using tools (like Collibra or Informatica) to track where data comes from and where it's going.
  • Change Management, Experience managing the technical "shock" to the business when switching from legacy BI tools to modern query engines like Starburst.

Nice to have

  • Bachelor's degree/University degree or equivalent experience Master's degree preferred ------------------------------------------------------

Skills

  • Our automated processing and AI do not involve relying on automatic or autonomous decision-making.
  • Please refer to any Jurisdictional Considerations, with specific provisions for your country (where relevant) for further details.
  • View Citi's EEO Policy Statement and the Know Your Rights poster.
  • Evaluate and integrate emerging tools and frameworks (e.g., Spark, Flink, Kafka) into the existing stack.

Compensation

  • $113,840.00 - $170,760.00 In addition to salary, Citi's offerings may also include, for eligible employees, discretionary and formulaic incentive and retention awards.

Benefits

  • Monitor system health and troubleshoot performance bottlenecks across the data lifecycle.

This listing is sourced directly from Citi's careers page and normalized into a canonical job model.