Role Summary
We are seeking an experienced
Data Engineer with strong expertise in
Databricks and modern
Lakehouse architecture to design, develop, and optimize scalable data pipelines. The ideal candidate will have hands-on experience with
Apache Spark (PySpark), Delta Lake, Python, SQL, and GCP, along with a strong understanding of data engineering best practices, performance optimization, and enterprise-scale data platforms.
Key Responsibilities
- Design, develop, and maintain scalable batch and streaming data pipelines using Databricks.
- Build robust ETL/ELT data pipelines using PySpark and Spark SQL.
- Develop and manage Delta Lake tables, ensuring ACID compliance, schema evolution, and efficient data management.
- Implement and maintain Lakehouse Medallion Architecture (Bronze, Silver, Gold).
- Perform data ingestion from multiple sources, including APIs, Kafka, cloud storage, and relational databases.
- Design and develop data models to support analytics, reporting, and business intelligence.
- Ensure data quality through validation, reconciliation, and monitoring processes.
- Optimize Spark workloads using partitioning, caching, query tuning, and cluster optimization techniques.
- Implement data governance and security using Unity Catalog or equivalent governance tools.
- Integrate data pipelines with orchestration tools such as Apache Airflow or Cloud Composer.
- Collaborate with business stakeholders, architects, and cross-functional teams to understand requirements and deliver scalable solutions.
- Troubleshoot production issues and continuously improve pipeline performance and reliability.
Mandatory SkillsDatabricks & Big Data
- Strong hands-on experience with Databricks (GCP preferred).
- Expertise in Apache Spark (PySpark).
- Experience with Delta Lake, including MERGE, UPSERT, Time Travel, and schema evolution.
- Strong understanding of Lakehouse Architecture.
Programming
- Strong proficiency in Python.
- Advanced SQL skills.
Data Engineering
- ETL/ELT pipeline development.
- Batch and streaming data processing.
- Data modeling (Dimensional Modeling/Lakehouse).
Cloud Platform- Hands-on experience with Google Cloud Platform (GCP).
- Experience with services such as:
- BigQuery
- Cloud Storage
- Cloud Composer
- Dataflow (preferred)
Good To Have Skills
- Apache Airflow / Cloud Composer.
- Kafka or other real-time streaming platforms.
- CI/CD using Git, GitHub Actions, or Azure DevOps.
- Infrastructure as Code (Terraform).
- Docker and Kubernetes.
- Exposure to BI tools such as Power BI or Tableau.
- Experience working in BFSI or other enterprise data environments.
Skills: databricks,gcp,sql,apache spark,delta lake