Abstract <p>To increase representativeness in data analysis, there is a need to extract and integrate data from various sources. This paper examines the application of machine learning to one of the main stages of data integration that is schema matching. A schema is the structure of a specific database, including the definition of relations, their attributes data types, and the definition of primary and foreign keys. The result of this stage is a matching of the elements of the source schemas and the elements of the target schema. A neural network model based on a combination of long short-term memory networks, attention mechanisms, and a multilayer perceptron is proposed. The model is trained using a hybrid set of features, including name similarity metrics, data types, tags, descriptive statistics, and correlation coefficients for numerical data. Experiments have been conducted showing that the proposed model outperforms the basic neural network model and classical schema matching methods. It is also shown that with federated learning, which preserves data privacy, the quality of the model is almost the same comparing to centralized learning.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Schema Matching Using Federated Learning on a Hybrid Feature Set

  • Andrey Vladimirovich Shepelev,
  • Sergey Alexandrovich Stupnikov

摘要

Abstract

To increase representativeness in data analysis, there is a need to extract and integrate data from various sources. This paper examines the application of machine learning to one of the main stages of data integration that is schema matching. A schema is the structure of a specific database, including the definition of relations, their attributes data types, and the definition of primary and foreign keys. The result of this stage is a matching of the elements of the source schemas and the elements of the target schema. A neural network model based on a combination of long short-term memory networks, attention mechanisms, and a multilayer perceptron is proposed. The model is trained using a hybrid set of features, including name similarity metrics, data types, tags, descriptive statistics, and correlation coefficients for numerical data. Experiments have been conducted showing that the proposed model outperforms the basic neural network model and classical schema matching methods. It is also shown that with federated learning, which preserves data privacy, the quality of the model is almost the same comparing to centralized learning.