Pyspark Bigdata Developer
Citi · Pune · 4+ yrs experience · Posted 2026-07-18
Tech stack: Python, SQL, Scala
Apply on the company site · Get a referral for this role
Citi salary & ratings · Citi interview process · More live openings
About the role
We are seeking a motivated and data-savvy Pyspark Bigdata Developer r to join our growing team. This role is perfect for an individual passionate about data, with foundational skills in both software development and quality assurance. The successful candidate will be responsible for developing and testing data-centric applications, with a strong focus on writing complex Strong programming skills in PySpark and Python in a BigData environment. This hybrid role offers a unique opportunity to work across the entire data lifecycle, from development and testing of data pipelines to delivering actionable insights through BI tools.
Responsibilities: - Strong programming skills in PySpark in a BigData environment.
- Familiarity with big data processing tools and techniques.
Qualifications: - with the Hadoop ecosystem including Hive, HDFS, Sqoop, Spark, Impala, Scala, etc.,Well versed with shell scripting and Autosys scheduler.
- Good understanding of distributed systems.
- Should be familiar with data warehouse concepts.
- with streaming data platforms.
- Excellent analytical and problem-solving skills.
- with writing complex SQL queries.
- on data modeling and data design is essential.
- Should be independent and resourceful dealing with risks/issues and resolving them in a timely manner.
- Excellent communication and articulation skills.
- Pyspark Strong programming skills in PySpark in a BigData environment.
- SQL: Strong proficiency in writing and optimizing complex SQL queries.
- Experience with a major relational database system (e.g., SQL Server, Hive, Impala ETL Concepts: Basic understanding of ETL (Extract, Transform, Load) processes and data pipeline concepts.
- Testing: Familiarity with data quality and testing methodologies.
- Version Control: Experience with version control systems, such as Git.
- Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent practical experience.
- Minimum 4 years of experience in a role involving data development, database management, or software quality assurance.
- Strong understanding of relational databases and data warehousing concepts.
- Foundational knowledge of at least one programming language (SQL,Pyspark,Python)
- Bachelor’s degree/University degree or equivalent experience
- This job description provides a high-level review of the types of work performed.
- Other job-related duties may be assigned as required.
Qualifications
- with the Hadoop ecosystem including Hive, HDFS, Sqoop, Spark, Impala, Scala, etc.,Well versed with shell scripting and Autosys scheduler.
- Good understanding of distributed systems.
- Should be familiar with data warehouse concepts.
- with streaming data platforms.
- Excellent analytical and problem-solving skills.
- with writing complex SQL queries.
- on data modeling and data design is essential.
- Should be independent and resourceful dealing with risks/issues and resolving them in a timely manner.
- Excellent communication and articulation skills.
- Pyspark Strong programming skills in PySpark in a BigData environment.
- SQL: Strong proficiency in writing and optimizing complex SQL queries.
- Experience with a major relational database system (e.g., SQL Server, Hive, Impala ETL Concepts: Basic understanding of ETL (Extract, Transform, Load) processes and data pipeline concepts.
- Testing: Familiarity with data quality and testing methodologies.
- Version Control: Experience with version control systems, such as Git.
- Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent practical experience.
- Minimum 4 years of experience in a role involving data development, database management, or software quality assurance.
- Strong understanding of relational databases and data warehousing concepts.
- Foundational knowledge of at least one programming language (SQL,Pyspark,Python)
- Bachelor’s degree/University degree or equivalent experience
- This job description provides a high-level review of the types of work performed.
- Other job-related duties may be assigned as required.
Responsibilities
- Strong programming skills in PySpark in a BigData environment.
- Familiarity with big data processing tools and techniques.
More openings at Citi
- IVR Testing Senior Analyst — Chennai
- Engineering Excellence Automation Testing Developer — Chennai
- Senior Pega Developer- Assistant Vice President — Chennai
- Applications Development Sr Programmer Analyst - C12 - CHENNAI — Chennai
- Java Microservice Application Development Manager — Chennai
- Assistant Vice President - Application Development — Chennai