Senior Big Data Engineer
Citi · Pune · 7+ yrs experience · Posted 2026-07-23
Tech stack: SQL, Scala
Apply on the company site · Get a referral for this role
Citi salary & ratings · Citi interview process · More live openings
About the role
Citi is looking for a Senior Big Data Engineer to design, build, and optimize large-scale data pipelines and distributed data systems that power critical business intelligence across the organization. Based in Pune and operating in a hybrid model, you will work within a high-performing engineering team where your expertise in PySpark, the Hadoop ecosystem, and streaming data platforms will directly shape the reliability and performance of Citi's data infrastructure.
Responsibilities: - Build and maintain scalable data pipelines using PySpark within a Big Data environment to process and transform large volumes of structured and unstructured data.
- Design and develop solutions across the Hadoop ecosystem — including Hive, HDFS, Sqoop, Spark, Impala, and Scala — to enable efficient data ingestion, processing, and storage.
- Develop and manage real-time and batch data workflows using streaming data platforms, ensuring high availability and low-latency data delivery.
- Write complex SQL queries to extract, validate, and analyze data across distributed systems, supporting data-driven decision-making.
- Design and implement data models and data architecture patterns aligned with data warehouse principles, ensuring scalability, accuracy, and consistency.
- Automate pipeline scheduling and orchestration using shell scripting and Autosys, reducing manual intervention and improving operational reliability.
- Independently identify, assess, and resolve technical risks and data issues in a timely manner, maintaining system integrity across the data platform.
Qualifications: - 7 years of relevant experience.
- Hands-on expertise in PySpark and Big Data processing, with the ability to build and optimize distributed data workflows at scale.
- Practical knowledge of the Hadoop ecosystem, including Hive, HDFS, Sqoop, Spark, Impala, and Scala, applied in a production environment.
- Proficiency in complex SQL query development for data analysis, transformation, and validation across large datasets.
- Solid understanding of distributed systems architecture and how data flows across interconnected processing layers.
- Demonstrated knowledge of data modelling and data design, with familiarity in data warehouse concepts and dimensional modelling techniques.
- Competence in shell scripting and job scheduling using Autosys or equivalent workflow automation tools.
- Strong analytical and problem-solving ability, with a track record of working independently to diagnose and resolve complex data engineering challenges.
- Clear and effective communication skills, with the ability to articulate technical concepts to both technical and non-technical audiences.
Qualifications
- 7 years of relevant experience.
- Hands-on expertise in PySpark and Big Data processing, with the ability to build and optimize distributed data workflows at scale.
- Practical knowledge of the Hadoop ecosystem, including Hive, HDFS, Sqoop, Spark, Impala, and Scala, applied in a production environment.
- Proficiency in complex SQL query development for data analysis, transformation, and validation across large datasets.
- Solid understanding of distributed systems architecture and how data flows across interconnected processing layers.
- Demonstrated knowledge of data modelling and data design, with familiarity in data warehouse concepts and dimensional modelling techniques.
- Competence in shell scripting and job scheduling using Autosys or equivalent workflow automation tools.
- Strong analytical and problem-solving ability, with a track record of working independently to diagnose and resolve complex data engineering challenges.
- Clear and effective communication skills, with the ability to articulate technical concepts to both technical and non-technical audiences.
Responsibilities
- Build and maintain scalable data pipelines using PySpark within a Big Data environment to process and transform large volumes of structured and unstructured data.
- Design and develop solutions across the Hadoop ecosystem — including Hive, HDFS, Sqoop, Spark, Impala, and Scala — to enable efficient data ingestion, processing, and storage.
- Develop and manage real-time and batch data workflows using streaming data platforms, ensuring high availability and low-latency data delivery.
- Write complex SQL queries to extract, validate, and analyze data across distributed systems, supporting data-driven decision-making.
- Design and implement data models and data architecture patterns aligned with data warehouse principles, ensuring scalability, accuracy, and consistency.
- Automate pipeline scheduling and orchestration using shell scripting and Autosys, reducing manual intervention and improving operational reliability.
- Independently identify, assess, and resolve technical risks and data issues in a timely manner, maintaining system integrity across the data platform.
More openings at Citi
- Senior Java Software Engineer Vice President — Pune
- GenAI intelligent insights, Conversational systems Lead - AVP — Pune
- IT Project Senior Analyst - Assistant Vice President — Chennai
- Data Analytics Senior Analyst - Assistant Vice President — Chennai
- Application Support Intermediate Analyst — Chennai
- Application Support Intermediate Analyst — Chennai