We are looking for an experienced Senior Data Engineer with strong expertise in PySpark, Python, and Informatica BDM development to support large-scale Data & Analytics initiatives within an enterprise banking environment. The ideal candidate should possess hands-on experience in building and maintaining robust ETL pipelines, data marts, and scalable data processing solutions across structured, semi-structured, and unstructured data sources.
The role requires strong analytical capabilities, end-to-end SDLC experience, and the ability to work closely with cross-functional teams throughout development, testing, deployment, and production support activities.
Requirements
Key Responsibilities
- Design, develop, and maintain scalable ETL pipelines and Data Mart solutions
- Develop high-performance data processing solutions using PySpark and Python
- Perform end-to-end SDLC activities including development, UAT support, bug fixing, production deployment, and post-production support
- Build and optimize data transformation pipelines for large-scale enterprise data platforms
- Perform data analysis, code debugging, and performance tuning across PySpark and SQL-based solutions
- Collaborate with business, analytics, and engineering teams to understand and implement data requirements
- Ensure data quality, integrity, scalability, and reliability across data pipelines
- Work with structured, semi-structured, and unstructured datasets in enterprise environments
- Participate in code reviews and implement software engineering best practices
- Support CI/CD implementation and data pipeline deployment activities
- Troubleshoot production issues and implement effective resolutions
- Contribute to technical documentation and knowledge-sharing initiatives
Required Technical Skills
Programming & Big Data Technologies
- Python
- PySpark
- Apache Spark
- Informatica BDM (Big Data Management)
- Hadoop
- MapReduce
- Hive
- Pandas
Data Engineering & ETL
- ETL Pipeline Development
- Data Mart Development
- Data Warehousing Concepts
- Data Transformation & Processing
- Data Pipeline Optimization
Databases & Query Languages
- Oracle SQL
- SQL
- NoSQL Databases
- Strong analytical and query-writing skills
Tools & Platforms
- Jupyter Notebook
- Git / Version Control
- CI/CD Pipelines
- Testing & Validation Frameworks
Required Experience
- Minimum 5+ years of commercial experience in Data Engineering or Data Analytics projects
- Strong hands-on experience in PySpark and Python-based ETL development
- Experience building enterprise-grade ETL pipelines and Data Mart solutions
- Strong experience in Informatica BDM development
- Experience handling end-to-end SDLC activities including development, UAT, production deployment, and post-production support
- Strong expertise in Oracle SQL and data analysis
- Hands-on experience debugging PySpark code and optimizing data processing workflows
- Experience working with production-grade data pipelines and large datasets
- Strong understanding of software engineering principles and coding best practices
- Experience working with Agile delivery environments
Preferred Domain Experience
- Banking
- Financial Services
- Digital Products
- Data & Analytics Platforms
Nice to Have
- Experience working with enterprise data lake and big data ecosystems
- Exposure to cloud-based data platforms
- Experience working with CI/CD and automated data pipeline deployments
- Knowledge of modern data engineering best practices and scalable architectures
Daily Tech Stack
The selected candidate will work extensively with:
- Python
- PySpark
- Informatica BDM
- Apache Spark
- Jupyter Notebook
- Oracle SQL
- SQL & NoSQL Databases
- Hadoop Ecosystem (Hive, MapReduce)
- CI/CD Tools
- ETL & Data Warehousing Technologies
Functional Competencies
- Strong problem-solving and analytical skills
- Excellent debugging and troubleshooting capabilities
- Ability to work independently in a fast-paced Agile environment
- Strong ownership mindset and attention to detail
- Effective stakeholder communication and collaboration skills
- Ability to manage multiple priorities and production support activities