<p>Data enhancement refers to the process of refining and optimizing datasets to improve the performance and generalizability of classification models. Traditionally, researchers have focused on methods like data balancing and feature selection to enhance classification results. While these approaches can improve model accuracy, their impact on the generalizability of classifiers remains limited. In response to this challenge, this study proposes a novel data enhancement method specifically designed to improve generalizability, a factor often overlooked in previous works. The method, referred to as “Find and Merge,” identifies instances within each class that exhibit high similarity, which can negatively affect the generalizability of the model. Using cosine similarity, the method groups similar instances, testing thresholds ranging from 0.75 to 0.95 to find the most optimal similarity ratio for improved generalizability. Once similar instances are identified, they are merged into a single representative instance using six different merging techniques: mean, median, inter quartile range (iqr), geometric mean, trimmed mean, and Tukey’s biweight. This merging process creates a more generalizable dataset by reducing redundancy while maintaining relevant patterns within the data. The proposed method is applied to several Coronary Artery Disease (CAD) datasets, including Z-Alizadeh Sani, Cleveland, Statlog, Heart Disease (HDS) and Firmingham (FHS) dataset. To evaluate the effectiveness of the enhanced datasets, classifiers such as Support Vector Machine (SVM), K-Nearest Neighbor (KNN), Random Forest (RF), and Ensemble Learning (EL) are used to assess performance across various metrics, including accuracy, precision, recall, specificity, F-score, and processing time. Both the original and enhanced datasets are compared to measure the impact of the “Find and Merge” method. The results demonstrate that the proposed approach significantly improves the generalizability of classifiers, offering a scalable solution that is not limited to a single dataset. This study’s key contribution lies in introducing a novel data enhancement strategy that goes beyond traditional balancing and feature selection methods, focusing directly on improving the generalization capacity of classification models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A new data analysis technique for sample reduction in healthcare datasets

  • Samet Aymaz

摘要

Data enhancement refers to the process of refining and optimizing datasets to improve the performance and generalizability of classification models. Traditionally, researchers have focused on methods like data balancing and feature selection to enhance classification results. While these approaches can improve model accuracy, their impact on the generalizability of classifiers remains limited. In response to this challenge, this study proposes a novel data enhancement method specifically designed to improve generalizability, a factor often overlooked in previous works. The method, referred to as “Find and Merge,” identifies instances within each class that exhibit high similarity, which can negatively affect the generalizability of the model. Using cosine similarity, the method groups similar instances, testing thresholds ranging from 0.75 to 0.95 to find the most optimal similarity ratio for improved generalizability. Once similar instances are identified, they are merged into a single representative instance using six different merging techniques: mean, median, inter quartile range (iqr), geometric mean, trimmed mean, and Tukey’s biweight. This merging process creates a more generalizable dataset by reducing redundancy while maintaining relevant patterns within the data. The proposed method is applied to several Coronary Artery Disease (CAD) datasets, including Z-Alizadeh Sani, Cleveland, Statlog, Heart Disease (HDS) and Firmingham (FHS) dataset. To evaluate the effectiveness of the enhanced datasets, classifiers such as Support Vector Machine (SVM), K-Nearest Neighbor (KNN), Random Forest (RF), and Ensemble Learning (EL) are used to assess performance across various metrics, including accuracy, precision, recall, specificity, F-score, and processing time. Both the original and enhanced datasets are compared to measure the impact of the “Find and Merge” method. The results demonstrate that the proposed approach significantly improves the generalizability of classifiers, offering a scalable solution that is not limited to a single dataset. This study’s key contribution lies in introducing a novel data enhancement strategy that goes beyond traditional balancing and feature selection methods, focusing directly on improving the generalization capacity of classification models.