- Posted 6 days ago
- Over 200 applicants have applied
Job Description
We are looking for Data Engineer.
Key Responsibilities
▪ Architect high-throughput pipelines processing petabytes of text and multimodal data.
▪ Implement deduplication, quality filtering and benchmark decontamination protocols.
▪ Lead licensing negotiations for proprietary datasets.
▪ Manage external vendors for large-scale digitisation and crowdsourcing.
▪ Own data provenance and the audit trail behind every corpus entering a training run.
Essential Skills & Experience
▪ Experience in data engineering or data acquisition at scale.
▪ Proven ownership of large-scale data infrastructure — Spark, Ray, Airflow, Kafka or equivalent.
▪ Practical deduplication and data-quality work on very large corpora.
▪ Strong vendor management and commercial negotiation ability.
▪ Python at production depth.
Preferred
▪ Web crawling and corpus construction at petabyte scale.
▪ Experience negotiating dataset licensing agreements.
Bachelor Of Technology (B.Tech/B.E), Master in Computer Application (M.C.A), Master OF Business Administration (M.B.A), Master of Commerce (M.Com), Master in Computer Management (M.C.M)
More Info
Key Skills
Data Pipelines
Petabyte Scale
Ray
Data Licensing
Distributed Data Processing
