Databricks AWS Engineer - Vice President
Citi · Pune · 10+ yrs experience · Posted 2026-07-18
Tech stack: AWS
Apply on the company site · Get a referral for this role
Citi salary & ratings · Citi interview process · More live openings
About the role
We are looking for a highly skilled Senior Databricks Engineer to contribute to the engineering, modernization, and continuous evolution of data processing platform on Databricks on AWS. While supporting the transition from the legacy Cloudera Hadoop platform to Databricks on AWS, this role will continue to play a key part in enhancing performance, simplifying pipelines, and delivering new capabilities on the Databricks platform over the long term.
Responsibilities: - Platform Engineering & Modernization
- Refactor and modernize existing Spark pipelines to Databricks native architectures
- Eliminate legacy Hadoop dependencies and adopt cloud native AWS patterns
- Enhance and extend existing processing logic using optimized Spark (JavaSpark / PySpark) on Databricks
- Databricks Native Development
- Build and optimize solutions using Databricks features, including Delta Lake, Databricks Workflows for orchestration and Auto scaling and job clusters
- Design & Solution Engineering
- Contribute to low and mid level architecture and design
- Translate high level architecture into detailed technical designs
- Define data models, pipeline patterns, and reusable components
- Ensure solutions are scalable, maintainable, and production ready
- Performance Optimization & Simplification
- Analyze, improve Spark job performance and simplify complex or over engineered pipelines into standardized, efficient patterns
- Engineering Standards & Best Practices
- Follow and contribute to Databricks and Spark engineering standards
- Write clean, modular, and testable code
- Contribute to shared frameworks, reusable libraries, and quality standards
- Collaboration & Stakeholder Engagement
- Work closely with senior architects, platform teams, and DevOps engineers
- Provide technical inputs, troubleshooting support, and implementation guidance
- Participate in design discussions and technical decision making
- Testing & Quality Assurance
- Develop unit, integration, and data validation tests
- Support production releases and post deployment validation
Qualifications: - Core Technical Skills
- 10+ years in data engineering or distributed systems
- Strong expertise in Apache Spark (JavaSpark / PySpark), Databricks on AWS, and Delta Lake with AWS services and large‑scale distributed data processing
- Modernization & Optimization Experience modernizing or refactoring legacy data platforms into cloud‑based architectures
- Strong background in Spark performance tuning and large‑scale batch optimization Design Capability
- Ability to translate architecture into implementable designs
- Understanding of data modeling and pipeline orchestration patterns Behavioral Competencies
- Strong problem‑solving mindset for complex distributed systems
- Comfortable working in time‑bound, high‑impact environments
- Proactive, accountable, and collaborative
- Clear communication skills across global teams
- Bachelor’s degree/University degree or equivalent experience
Qualifications
- Core Technical Skills
- 10+ years in data engineering or distributed systems
- Strong expertise in Apache Spark (JavaSpark / PySpark), Databricks on AWS, and Delta Lake with AWS services and large‑scale distributed data processing
- Modernization & Optimization Experience modernizing or refactoring legacy data platforms into cloud‑based architectures
- Strong background in Spark performance tuning and large‑scale batch optimization Design Capability
- Ability to translate architecture into implementable designs
- Understanding of data modeling and pipeline orchestration patterns Behavioral Competencies
- Strong problem‑solving mindset for complex distributed systems
- Comfortable working in time‑bound, high‑impact environments
- Proactive, accountable, and collaborative
- Clear communication skills across global teams
- Bachelor’s degree/University degree or equivalent experience
Responsibilities
- Platform Engineering & Modernization
- Refactor and modernize existing Spark pipelines to Databricks native architectures
- Eliminate legacy Hadoop dependencies and adopt cloud native AWS patterns
- Enhance and extend existing processing logic using optimized Spark (JavaSpark / PySpark) on Databricks
- Databricks Native Development
- Build and optimize solutions using Databricks features, including Delta Lake, Databricks Workflows for orchestration and Auto scaling and job clusters
- Design & Solution Engineering
- Contribute to low and mid level architecture and design
- Translate high level architecture into detailed technical designs
- Define data models, pipeline patterns, and reusable components
- Ensure solutions are scalable, maintainable, and production ready
- Performance Optimization & Simplification
- Analyze, improve Spark job performance and simplify complex or over engineered pipelines into standardized, efficient patterns
- Engineering Standards & Best Practices
- Follow and contribute to Databricks and Spark engineering standards
- Write clean, modular, and testable code
- Contribute to shared frameworks, reusable libraries, and quality standards
- Collaboration & Stakeholder Engagement
- Work closely with senior architects, platform teams, and DevOps engineers
- Provide technical inputs, troubleshooting support, and implementation guidance
- Participate in design discussions and technical decision making
- Testing & Quality Assurance
- Develop unit, integration, and data validation tests
- Support production releases and post deployment validation