Augmented Data Warehouses for Value Capture
摘要
Nobody can deny the monumental impact of Big Data on Data Warehouse ( \(\mathcal{D}\mathcal{W}\) ) technology. This impact has been materialized by the spectacular adaptation of Big Data solutions, initially developed to deal with the different Big Data V’s (Volume, Variety, Velocity, Veracity), to the context of DW. This gives raise to the concept of augmented DW. The usage of MapReduce and Spark to manage volume, stream processing systems for deploying stream DW, and multimodel DW to handle data variety, are examples of these adaptations. By deeply analyzing these adaptations, we figure out that the V corresponding to Big Data Value did not get the same attention as the other Vs. DWs are built to capture value, otherwise, they will have a marginal utility. One of the major problems of the value capturing in DW concerns the lack of data when performing the warehouse exploration through OLAP queries. This situation is caused by the absence of concepts/hierarchies/instances satisfying these queries in the initial sources used to build the target DW. To increase the utility of a DW in terms of value capture, the selection of alternative resources is recommended. In this paper, we first illustrate the different scenarios requiring concept and data enrichment and focus on the usage of external resources such as Linked Open Data and ontologies to deal with the absence of data. Secondly, we propose a technique to value capturing by revisiting the \(\mathcal{D}\mathcal{W}\) life cycle phases. Value metrics are given based on the value-capturing requirements. Finally, experiments are conducted to show the effectiveness of our proposal.