PySpark Big Data Developer
Citi · Pune · 4–8 yrs experience · Posted 2026-07-18
Tech stack: Python, SQL
Apply on the company site · Get a referral for this role
Citi salary & ratings · Citi interview process · More live openings
About the role
We are seeking a highly skilled and experienced BigData/PySpark Engineer to join our dynamic Big Data Analytics team. This role is pivotal in designing, developing, and optimizing robust, scalable data pipelines for large-scale data processing and analytics.
Responsibilities: - Design & Development:
- Create and optimize scalable ETL (Extraction, Transformation, Loading) pipelines using PySpark for massive datasets.
- Coding & Engineering: Write clean, efficient, well-documented code primarily in Python (PySpark) often leveraging frameworks/tools.
- Collaboration: Work with cross-functional teams (senior developers, data engineers, analysts, business partners) to understand data requirements and ensure seamless solution integration.
- Troubleshooting & Optimization: Debug and resolve data processing issues and performance bottlenecks in Spark applications and other big data technologies.
- Full SDLC Involvement:
- Participate in the entire software development lifecycle, from requirements analysis and design to testing, deployment, and operations.
- Data Integrity: Ensure high data quality and integrity throughout the data lifecycle.
- This candidate possesses
- 4-8 years of experience in developing and managing Enterprise Applications, demonstrating a robust foundation in Big Data technologies and a strong grasp of software development principles.
- Key Experience & Expertise:
- Enterprise Application Development:
- 4-8 years in developing and managing enterprise-grade applications.
- Object-Oriented Programming (OOP):
- Solid foundation in OOP concepts.
- Big Data Development: Expertise in PySpark, HDFS, Hive, Sqoop, and Hadoop for Big Data environments.
- Database Technologies: Good exposure to SQL Server and ORACLE databases.
- Experience with query writing for data validation/manipulation
- Scripting & Automation: Proficient in Shell Scripting and experience with job scheduling tools like Autosys
- BI Reporting Tools:
- Some exposure to BI tools, specifically Tableau
- Tools & Practices:
- Proficient with Git; experience with JIRA, Confluence.
- Familiarity with DevOps and CI/CD pipelines.
Qualifications: - requirements and ensure seamless solution integration.
- Troubleshooting & Optimization: Debug and resolve data processing issues and performance bottlenecks in Spark applications and other big data technologies.
- Full SDLC Involvement:
- Participate in the entire software development lifecycle, from requirements analysis and design to testing, deployment, and operations.
- Data Integrity: Ensure high data quality and integrity throughout the data lifecycle.
- This candidate possesses
- 4-8 years of experience in developing and managing Enterprise Applications, demonstrating a robust foundation in Big Data technologies and a strong grasp of software development principles.
- Key Experience & Expertise:
- Enterprise Application Development:
- 4-8 years in developing and managing enterprise-grade applications.
- Object-Oriented Programming (OOP):
- Solid foundation in OOP concepts.
- Big Data Development: Expertise in PySpark, HDFS, Hive, Sqoop, and Hadoop for Big Data environments.
- Database Technologies: Good exposure to SQL Server and ORACLE databases.
- Experience with query writing for data validation/manipulation
- Scripting & Automation: Proficient in Shell Scripting and experience with job scheduling tools like Autosys
- BI Reporting Tools:
- Some exposure to BI tools, specifically Tableau
- Tools & Practices:
- Proficient with Git; experience with JIRA, Confluence.
- Familiarity with DevOps and CI/CD pipelines.
Qualifications
- requirements and ensure seamless solution integration.
- Troubleshooting & Optimization: Debug and resolve data processing issues and performance bottlenecks in Spark applications and other big data technologies.
- Full SDLC Involvement:
- Participate in the entire software development lifecycle, from requirements analysis and design to testing, deployment, and operations.
- Data Integrity: Ensure high data quality and integrity throughout the data lifecycle.
- This candidate possesses
- 4-8 years of experience in developing and managing Enterprise Applications, demonstrating a robust foundation in Big Data technologies and a strong grasp of software development principles.
- Key Experience & Expertise:
- Enterprise Application Development:
- 4-8 years in developing and managing enterprise-grade applications.
- Object-Oriented Programming (OOP):
- Solid foundation in OOP concepts.
- Big Data Development: Expertise in PySpark, HDFS, Hive, Sqoop, and Hadoop for Big Data environments.
- Database Technologies: Good exposure to SQL Server and ORACLE databases.
- Experience with query writing for data validation/manipulation
- Scripting & Automation: Proficient in Shell Scripting and experience with job scheduling tools like Autosys
- BI Reporting Tools:
- Some exposure to BI tools, specifically Tableau
- Tools & Practices:
- Proficient with Git; experience with JIRA, Confluence.
- Familiarity with DevOps and CI/CD pipelines.
Responsibilities
- Design & Development:
- Create and optimize scalable ETL (Extraction, Transformation, Loading) pipelines using PySpark for massive datasets.
- Coding & Engineering: Write clean, efficient, well-documented code primarily in Python (PySpark) often leveraging frameworks/tools.
- Collaboration: Work with cross-functional teams (senior developers, data engineers, analysts, business partners) to understand data requirements and ensure seamless solution integration.
- Troubleshooting & Optimization: Debug and resolve data processing issues and performance bottlenecks in Spark applications and other big data technologies.
- Full SDLC Involvement:
- Participate in the entire software development lifecycle, from requirements analysis and design to testing, deployment, and operations.
- Data Integrity: Ensure high data quality and integrity throughout the data lifecycle.
- This candidate possesses
- 4-8 years of experience in developing and managing Enterprise Applications, demonstrating a robust foundation in Big Data technologies and a strong grasp of software development principles.
- Key Experience & Expertise:
- Enterprise Application Development:
- 4-8 years in developing and managing enterprise-grade applications.
- Object-Oriented Programming (OOP):
- Solid foundation in OOP concepts.
- Big Data Development: Expertise in PySpark, HDFS, Hive, Sqoop, and Hadoop for Big Data environments.
- Database Technologies: Good exposure to SQL Server and ORACLE databases.
- Experience with query writing for data validation/manipulation
- Scripting & Automation: Proficient in Shell Scripting and experience with job scheduling tools like Autosys
- BI Reporting Tools:
- Some exposure to BI tools, specifically Tableau
- Tools & Practices:
- Proficient with Git; experience with JIRA, Confluence.
- Familiarity with DevOps and CI/CD pipelines.
More openings at Citi
- Assistant Vice President Data Science and Gen AI — Gurugram
- Quantitative Analyst and Developer — Mumbai
- DevOps Senior Programmer - Assistant Vice President — Pune
- Java Fullstack Developer with React and J2EE — Pune
- Assistant Vice President - Application Development — Pune
- Senior Java Developer - Assistant Vice President — Pune