错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Introduction

  • Zhixin Qi,
  • Hongzhi Wang,
  • Zejiao Dong

摘要

In recent years, with the development of the information age, the amount of data has grown dramatically. At the same time, dirty data have already existed in various types of databases. Due to the negative impacts of dirty data on data mining and machine learning results, data quality issues have attracted widespread attention. Motivated by this, this book aims to analyze the impacts of dirty data on machine learning models and explore the proper methods for dirty data processing. This chapter discusses the background of dirty data processing for machine learning. In Sect. 1.1, we analyze three basic dimensions of data quality to motivate the necessity of processing dirty data in the database and machine learning communities. In Sect. 1.2, we summarize the existing studies and explain the differences of our research and current work. We conclude the chapter with an overview of the structure of this book in Sect. 1.3.