Machine learning paradigms are evolved as general need algorithms for data science applications. Unfortunately, their performance is degraded when the quality of the dataset is not up to the mark. It does mean that the quality of inputs have an impact on the outputs. Therefore, it became an essential paradigm to have feature selection mechanisms prior to applying machine learning techniques. An efficient classifier may perform poorly or fail in performing its intended work when the data has redundant and irrelevant features. Many approaches came into existence for feature selection. They include embedded methods, filter methods and wrapper methods. The current article proposed a feature extraction approach known as Weighted Normalized Mutual Information (WNII) which exploits both filter and wrapper mechanisms. This approach works on given samples and implicitly gets rid of bias and redundancies from the given dataset. A pragmatic study is made with different datasets and the results evaluated with many standard approaches. The proposed method is found to have better performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Normalized Mutual Information-Driven Feature Extraction Method for Big Data Analytics

  • Raghuram Bhukya

摘要

Machine learning paradigms are evolved as general need algorithms for data science applications. Unfortunately, their performance is degraded when the quality of the dataset is not up to the mark. It does mean that the quality of inputs have an impact on the outputs. Therefore, it became an essential paradigm to have feature selection mechanisms prior to applying machine learning techniques. An efficient classifier may perform poorly or fail in performing its intended work when the data has redundant and irrelevant features. Many approaches came into existence for feature selection. They include embedded methods, filter methods and wrapper methods. The current article proposed a feature extraction approach known as Weighted Normalized Mutual Information (WNII) which exploits both filter and wrapper mechanisms. This approach works on given samples and implicitly gets rid of bias and redundancies from the given dataset. A pragmatic study is made with different datasets and the results evaluated with many standard approaches. The proposed method is found to have better performance.