Job Title: Data Engineering + Data Modelling Lead
Experience: 6–8 Years
Location: Bangalore (Preferred) / Remote
Job Responsibilities:
- Lead end-to-end data engineering and data modeling initiatives, including requirements gathering, solution design, development, testing, deployment, andoperational support.
- Develop scalable, high-performance data pipelines using PySpark and Apache Spark for data ingestion, transformation, integration, and loading acrossenterprise platforms.
- Define and implement enterprise-wide conceptual, logical, and physical data models to support business, analytical, and operational requirements.
- Design and maintain database schemas, tables, views, indexes, and data structures that support both transactional and analytical workloads.
- Lead the development of ETL/ELT frameworks and reusable components to ensure efficient and standardized data processing.
- Implement data quality frameworks, validation rules, profiling techniques, and monitoring processes to ensure data accuracy, completeness, consistency, andreliability.
- Optimize data models, ETL pipelines, database performance, and Spark workloads to improve scalability, processing efficiency, and query performance.
- Document data models, data lineage, ETL workflows, metadata, and technical designs in a clear and comprehensive manner.
- Provide technical leadership and mentorship to development teams.
Required Experience & Skills:
- 6-8 years of experience in Data Engineering, ETL/ELT Development, and Data Modeling, with demonstrated experience leading enterprise-scale data initiatives and deliveryteams.
- Strong expertise in Databricks/Snowflake(or similar ETL tools), PySpark, Apache Spark, Spark SQL, and Python for designing, developing, and optimizing high-volume datapipelines and distributed data processing solutions.
- Proven experience in designing Conceptual, Logical, and Physical Data Models, including ER Modeling, Dimensional Modeling, Data Vault, and normalization/denormalizationtechniques.
- Hands-on experience with data lake, data warehouse, and cloud-based data platforms such as Databricks, Snowflake, AWS, or Azure.
- Strong knowledge of relational databases and performance tuning, including Oracle, SQL Server, PostgreSQL, Snowflake, schema design, indexing, and query optimization.
- Experience with data integration and orchestration technologies such as Apache Airflow, Kafka, NiFi, or equivalent enterprise data movement tools.
- Solid understanding of data governance, metadata management, data lineage, and data quality frameworks, with the ability to establish enterprise data standards and bestpractices.