Data Preprocessing and Feature Selection Approach for an Automated Expert Finding System for Academic Events
摘要
The recent developments in the data science and availability of huge operational data proved to be promising. It extended better possibilities for developing data-driven intelligent expert systems. While exploring such vast datasets, data pre-preprocessing is an essential phase with the goal to produce data which can be granted as correct and precise. Such preprocessed dataset is useful for applying further data mining algorithms. On the other side, such a good quality data can produce outstanding results using a simple algorithm too. Data collected from real world mostly contains impurities like noise, multiple delimiters, extraneous characters, missing values, inconsistency and may not be in useful format. Data preprocessing is required task for cleaning such raw data. It is imperative step as a part of effective data evaluation as it eliminates inconsistencies and duplications in data and converts from one format to another expected format. It pertains to a collection of techniques aimed at improving data quality by applying cleaning, reduction and transformation. Feature selection considers selecting a subset or list of attributes or variables that are used to apply mining algorithms and as well describe that data accurately. Feature selection is having significance in domains where features are many and comparatively samples (data points) are only few. The task of automatically finding expert resource persons in any field like academic, medical, law, finance and many other such fields is challenging as there is availability of large-scale expertise-related information and is generated from various data resources. This article discusses the data cleaning and feature selection approach applied for an automated expert finding system for academic events. Approximately, 10% raw records are dropped as a result of data cleaning.