Systems Ltd is looking for a Forward Deployed Engineer – Data Management to Make enterprise data and knowledge usable by AI — build AI-ready data products, knowledge graphs, and retrieval infrastructure that every other practice depends on.
KEY RESPONSIBILITIES
- Build AI-ready data pipelines and data products that other practices can consume directly
- Design and build knowledge graphs and semantic layers that structure enterprise knowledge for AI consumption
- Own vector and retrieval infrastructure (embeddings, indexes, hybrid search) shared across GenAI and ML practices
- Run data quality assessment and remediation specifically for AI/ML consumption, not just BI
- Own knowledge engineering — taxonomy, ontology, and ingestion pipelines for enterprise knowledge sources
- Partner with GenAI Engineers, Data Scientists, and AI Architects to expose curated data/knowledge as reusable assets
- Explain the difference between BI-grade and AI-grade data quality to non-technical stakeholders
- Act as a shared upstream dependency for multiple practices — manage competing requests
- Document data/knowledge assets clearly enough for self-service reuse
REQUIREMENTS & SKILLS
- 5–10+ yrs data engineering, with 2+ yrs building AI-ready data products specifically
- Strong in knowledge graph technologies (Neo4j, RDF/SPARQL, or similar) and semantic/ontology modeling
- Experience with vector/retrieval infrastructure (embeddings, ANN indexes, hybrid search)
- Solid data pipeline engineering (Spark, dbt, Airflow, or similar) and data quality frameworks
- Familiarity with enterprise data governance and lineage tooling
- Can explain the difference between BI-grade and AI-grade data quality to a non-technical stakeholder
- Collaborates closely with GenAI Engineers, Data Scientists, and AI Architects as a shared upstream dependency
- Prioritizes competing requests from multiple practices fairly and transparently
- Documents clearly enough that other teams can self-serve without hand-holding
- Success metrics: data/knowledge asset reuse across practices · data quality incidents affecting AI systems (target zero) · time from raw data to AI-ready asset