Strong Middle Data Engineer – Customer Platform & Data
Position Overview
We are seeking Strong Middle Data Engineers with solid experience in distributed data systems, Spark-based data processing, and production-grade data platform engineering.
In this role, you will build, scale, maintain, and improve large-scale data platforms supporting critical initiatives across customer data, data pipeline expansion, compliance-related engineering, new data source integration, and platform modernization.
The ideal candidate is a hands-on engineer who can work independently, take ownership of technical solutions, contribute to system design and implementation, and operate effectively in a highly automated development environment.
The engineering organization is evolving toward a spec-led and agentic development model, where engineers increasingly focus on defining requirements, designing data models, creating technical specifications, reviewing AI-generated code, and ensuring overall solution quality.
Key Responsibilities
- Develop, maintain, and optimize batch and streaming data pipelines using Scala and Apache Spark.
- Build and support Kafka-based data processing and streaming solutions.
- Create and maintain Apache Airflow DAGs, workflows, and data processing schedules.
- Design and implement data models, schemas, mappings, and validation checks.
- Develop scalable ETL/ELT pipelines using SQL and distributed data processing technologies.
- Monitor pipeline health, investigate failures, and troubleshoot production issues.
- Support deployments and ensure reliable operation of production data platforms.
- Implement data quality and validation controls to ensure accurate and reliable data.
- Review and validate AI-generated code and proposed implementations.
- Collaborate with senior engineers, product teams, and data consumers to deliver technical solutions.
- Contribute to system design, technical specifications, and modernization initiatives.
- Take ownership of assigned systems and components, including reliability, maintainability, and performance.
Required Qualifications
- 3+ years of professional Data Engineering experience.
- Strong hands-on experience with Scala and Apache Spark.
- Experience working with batch processing and distributed data systems.
- Practical experience with Kafka, Kafka Streams, Flink, or similar streaming technologies.
- Hands-on experience with Apache Airflow, including DAG development and workflow management.
- Strong knowledge of SQL, data modeling, and ETL/ELT development.
- Experience with software engineering practices including:
- Automated testing
- Code reviews
- Git/version control
- CI/CD
- Production deployments
- Strong troubleshooting and production support capabilities.
- Ability to work independently and take ownership of technical solutions.
- Good communication and collaboration skills.
- Ability to understand technical requirements and translate them into reliable data engineering solutions.
Nice-to-Have Qualifications
- Hands-on Apache Flink experience.
- Experience with Cassandra, DynamoDB, ScyllaDB, or other NoSQL databases.
- Experience working with customer data, clickstream, loyalty, booking, or transactional data.
- Experience with GitHub Copilot, Claude, Cursor, or other AI-assisted development tools.
- Understanding of data quality monitoring, observability, and pipeline health monitoring.
- Experience with specification-driven or spec-to-code development.
- Familiarity with agentic development frameworks and engineering practices.
Success Profile
The successful candidate is a strong middle-level Data Engineer who can work with limited supervision while knowing when to involve senior engineers.
You should be able to demonstrate:
- Practical experience building and maintaining Spark-based data pipelines.
- Solid Scala development experience in production.
- Understanding of distributed data processing and batch workloads.
- Hands-on experience with Kafka or comparable streaming technologies.
- Ability to build and manage Airflow DAGs and production workflows.
- Strong SQL, ETL/ELT, and data modeling skills.
- Ability to troubleshoot pipeline failures and production data issues.
- Experience participating in code reviews, testing, CI/CD, and deployments.
- Ability to review AI-generated code critically and identify incorrect, inefficient, or unsafe implementations.
- An ownership mindset toward data quality, reliability, maintainability, and delivery.
- Ability to communicate effectively with senior engineers, product teams, and data consumers.
Engineering Environment
You will work in a modern, highly automated data engineering environment using technologies including:
Scala | Apache Spark | Kafka | Kafka Streams | Apache Flink | Apache Airflow | SQL | NoSQL | CI/CD | AI-Assisted Development
The team is moving toward a spec-led, agentic engineering model, giving engineers greater responsibility for requirements definition, technical specifications, data modeling, solution design, and validation of AI-assisted implementations.
Why This Position
- Work on large-scale data platforms supporting high-priority business initiatives.
- Gain exposure to complex customer, transactional, compliance, and operational data.
- Work with modern distributed data technologies including Spark, Scala, Kafka, and Airflow.
- Participate in the modernization and expansion of enterprise data platforms.
- Develop skills in AI-assisted and agentic software development.
- Collaborate closely with experienced senior engineers while owning meaningful technical components.
- Work in an engineering culture focused on automation, quality, reliability, and continuous improvement.