错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Balancing data imbalance in biomedical datasets using a stacked augmentation approach with STDA, DAGAN, and pufferfish optimization to reveal AI's transformative impact

  • Bhaskar Kumar Veedhi,
  • Kaberi Das,
  • Debahuti Mishra,
  • Sashikala Mishra,
  • Mandakini Priyadarshani Behera

摘要

This study systematically assesses the effectiveness of data augmentation techniques, focusing particularly on the stacked approach, across a diverse range of biomedical datasets. Advanced data augmentation methods, including style transfer data augmentation (STDA) and data augmentation generative adversarial network (DAGAN), offer promising solutions in this context. STDA leverages style transfer techniques from computer vision to create varied and realistic synthetic samples while preserving semantic content, thereby enhancing dataset variability and improving model generalization. Conversely, DAGAN utilizes generative adversarial networks (GANs) to learn the underlying data distribution, resulting in the generation of high-quality synthetic samples. Subsequently, STDA and DAGAN techniques are applied to the selected data, followed by convolution and pooling operations on the augmented samples. The pooled samples are then stacked, and the Cosine similarity is calculated to measure the resemblance to the original dataset. Parameters of STDA and DAGAN are fine-tuned to maximize the Cosine similarity. Additionally, a metaheuristic optimization algorithm called Pufferfish optimization algorithm (POA) is employed to optimize decision variables—learning rate, number of hidden units (common for both STDA and DAGAN), number of attention heads for STDA, and latent space dimensionality of DAGAN. The classification performance of the proposed POA-stacked STDA-DAGAN data augmentation methodology, along with variants using genetic algorithm (GA) and particle swarm optimization (PSO), is compared with original datasets and SMOTE balanced datasets using various classifiers and the observed accuracies, sensitivity, specificity and F-score are ranging from 91.83% to 99.11%, 91.38% to100%, 92.50% to 100% and 92.95% to 98.32% respective for all ten biomedical datasets used in this research. Additionally, the class distribution graphs, learning curves, statistical validation, and execution times are recorded to demonstrate the effectiveness of the proposed approach.