<p>Sparse numerical datasets are dominant in fields such as applied mathematics, astronomy, finance, and healthcare, presenting challenges due to their high dimensionality and sparse distribution. The predominance of zero values complicates optimal feature selection, making data analysis and model performance more complex. To overcome this challenge, this study introduces a deep learning-based algorithm, Hybrid Stacked Sparse Autoencoder (HSSAE), which integrates <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_20534_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="18" /> </InlineMediaObject> <EquationSource Format="TEX">\({\text{L}}_{1}\)</EquationSource> </InlineEquation> and <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_20534_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="18" /> </InlineMediaObject> <EquationSource Format="TEX">\({\text{L}}_{2}\)</EquationSource> </InlineEquation> regularization with binary cross-entropy loss to improve feature selection efficiency, where <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_20534_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="18" /> </InlineMediaObject> <EquationSource Format="TEX">\({\text{L}}_{1}\)</EquationSource> </InlineEquation> regularization penalizes large weights, simplifying data representations, while <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_20534_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="18" /> </InlineMediaObject> <EquationSource Format="TEX">\({\text{L}}_{2}\)</EquationSource> </InlineEquation> regularization prevents overfitting by limiting the total weight size. Additionally, the dropout technique enhances the algorithm’s performance by randomly deactivating neurons during training, avoiding over-reliance on specific features. Meanwhile, batch normalization stabilizes weight distributions, reducing computational complexity and accelerating the convergence. The proposed algorithm, HSSAE, was evaluated against traditional classifiers, including Decision Tree, Random Forest, K-Nearest Neighbors, and Naïve Bayes, as well as deep learning-based models, such as Convolutional Neural Network, Long Short-Term Memory, and Stacked Sparse Autoencoder, in terms of Precision, Recall, Accuracy, F1-score, AUC, and Hamming Loss. Quantitatively, the proposed algorithm, HSSAE, was tested on two different sparse datasets, demonstrating superior performance with the highest accuracy of 89% on the health indicator dataset and 93% on the EHRs diabetes prediction dataset, respectively, and outperforming competing classifiers. The proposed algorithm, HSSAE, extracts features effectively and enhances robustness, making it well-suited for sparse data applications, particularly in healthcare, where high prediction accuracy is crucial.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A deep learning framework with hybrid stacked sparse autoencoder for type 2 diabetes prediction

  • Abdussamad,
  • Hanita Daud,
  • Rajalingam Sokkalingam,
  • Muhammad Zubair,
  • Iliyas Karim Khan,
  • Zafar Mahmood

摘要

Sparse numerical datasets are dominant in fields such as applied mathematics, astronomy, finance, and healthcare, presenting challenges due to their high dimensionality and sparse distribution. The predominance of zero values complicates optimal feature selection, making data analysis and model performance more complex. To overcome this challenge, this study introduces a deep learning-based algorithm, Hybrid Stacked Sparse Autoencoder (HSSAE), which integrates \({\text{L}}_{1}\) and \({\text{L}}_{2}\) regularization with binary cross-entropy loss to improve feature selection efficiency, where \({\text{L}}_{1}\) regularization penalizes large weights, simplifying data representations, while \({\text{L}}_{2}\) regularization prevents overfitting by limiting the total weight size. Additionally, the dropout technique enhances the algorithm’s performance by randomly deactivating neurons during training, avoiding over-reliance on specific features. Meanwhile, batch normalization stabilizes weight distributions, reducing computational complexity and accelerating the convergence. The proposed algorithm, HSSAE, was evaluated against traditional classifiers, including Decision Tree, Random Forest, K-Nearest Neighbors, and Naïve Bayes, as well as deep learning-based models, such as Convolutional Neural Network, Long Short-Term Memory, and Stacked Sparse Autoencoder, in terms of Precision, Recall, Accuracy, F1-score, AUC, and Hamming Loss. Quantitatively, the proposed algorithm, HSSAE, was tested on two different sparse datasets, demonstrating superior performance with the highest accuracy of 89% on the health indicator dataset and 93% on the EHRs diabetes prediction dataset, respectively, and outperforming competing classifiers. The proposed algorithm, HSSAE, extracts features effectively and enhances robustness, making it well-suited for sparse data applications, particularly in healthcare, where high prediction accuracy is crucial.