Introduction
摘要
In recent years, with the development of the information age, the amount of data has grown dramatically. At the same time, dirty data have already existed in various types of databases. Due to the negative impacts of dirty data on data mining and machine learning results, data quality issues have attracted widespread attention. Motivated by this, this book aims to analyze the impacts of dirty data on machine learning models and explore the proper methods for dirty data processing. This chapter discusses the background of dirty data processing for machine learning. In Sect. 1.1, we analyze three basic dimensions of data quality to motivate the necessity of processing dirty data in the database and machine learning communities. In Sect. 1.2, we summarize the existing studies and explain the differences of our research and current work. We conclude the chapter with an overview of the structure of this book in Sect. 1.3.