About Relationships in Data Lakes
摘要
In the era of Big Data, managing voluminous and heterogeneous data presents significant challenges for organizations. To tackle these challenges, the concept of a data lake has emerged as a promising solution, allowing the storage of raw data from diverse sources in their original format. An efficient metadata management system plays a crucial role in preventing data lake to turn into an unusable data swamp by providing a structured framework for organizing, categorizing and establishing relationships between data entities. In this paper, identify the various relationships from diverse domains found in the literature. Then, we categorize the types of relationships and propose a relationship typology that classes relationships by similarity, containment, grouping and provenance. Eventually, we also aim to check whether goldMEDAL, a state-of-the-art generic metadata management model, adequately supports all such relationships. This evaluation is particularly relevant for Bial-X, which seeks to implement a robust metadata management system based on goldMEDAL’s concepts.