Data quality assessment is one of the most fundamental operations executed during data integration. Data validity is a collection of validation rules applied to the dataset’s attributes. The validation rules provided by domain experts must be known during the data validation checks. In practice domain experts may not be available, or their number is insufficient, and the project timeline may be strict, resulting in unknown data validity rules. In this paper, we present a low code framework for outlier and anomaly detection as an alternative to traditional data validation rules. The framework comprises statistical and machine learning methods that recognize outliers and label them as invalid data. We evaluated the accuracy and scalability of four methods using the TPC-DI benchmark tests, and the results indicate a high level of correctness when the appropriate method is employed for a specific distribution. The framework is exceptionally effective for data with skewed and normal distributions for local and global outliers. Additionally, for different scaling factors, we observed that the ratio of validation time to data integration execution time is low and stable as datasets increase.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Validation Without Rules: A Data Integration Case Study

  • Stefan Dzalev,
  • Goran Velinov

摘要

Data quality assessment is one of the most fundamental operations executed during data integration. Data validity is a collection of validation rules applied to the dataset’s attributes. The validation rules provided by domain experts must be known during the data validation checks. In practice domain experts may not be available, or their number is insufficient, and the project timeline may be strict, resulting in unknown data validity rules. In this paper, we present a low code framework for outlier and anomaly detection as an alternative to traditional data validation rules. The framework comprises statistical and machine learning methods that recognize outliers and label them as invalid data. We evaluated the accuracy and scalability of four methods using the TPC-DI benchmark tests, and the results indicate a high level of correctness when the appropriate method is employed for a specific distribution. The framework is exceptionally effective for data with skewed and normal distributions for local and global outliers. Additionally, for different scaling factors, we observed that the ratio of validation time to data integration execution time is low and stable as datasets increase.