<p>In this paper, we investigate optimization strategies for predicting high-mortality diseases in response to the challenges of unstable data quality and missing labels in healthcare big data. Firstly, a systematic review of the current state of research and application of machine learning algorithms for predicting heart disease, cardiovascular and cerebrovascular disease, lung cancer, diabetes, and breast cancer is presented. The core contributions are: 1) Semi-supervised self-training experiments were conducted on 14 datasets with labeling ratios of 30<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\%\)</EquationSource> </InlineEquation>-95<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\%\)</EquationSource> </InlineEquation>, and the performance of classifiers such as SVM, KNN, LR, AB, DT, and RF was evaluated, and it was found that data balance is crucial for the accuracy improvement of most of the classifiers, and that the sensitivities to labeling ratios varied significantly between different classifiers, e.g., SVM robustness and KNN dependency. 2) In-depth analysis of feature importance identifies the key attributes of each dataset, whose absence significantly degrades the model performance, while RF exhibits optimal robustness due to integration properties. The study provides methodological references and empirical evidence for precision medicine prediction under limited labeling conditions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A review of predictive healthcare treatment based on text mining

  • Mukun Cao,
  • Bing Li

摘要

In this paper, we investigate optimization strategies for predicting high-mortality diseases in response to the challenges of unstable data quality and missing labels in healthcare big data. Firstly, a systematic review of the current state of research and application of machine learning algorithms for predicting heart disease, cardiovascular and cerebrovascular disease, lung cancer, diabetes, and breast cancer is presented. The core contributions are: 1) Semi-supervised self-training experiments were conducted on 14 datasets with labeling ratios of 30 \(\%\) -95 \(\%\) , and the performance of classifiers such as SVM, KNN, LR, AB, DT, and RF was evaluated, and it was found that data balance is crucial for the accuracy improvement of most of the classifiers, and that the sensitivities to labeling ratios varied significantly between different classifiers, e.g., SVM robustness and KNN dependency. 2) In-depth analysis of feature importance identifies the key attributes of each dataset, whose absence significantly degrades the model performance, while RF exhibits optimal robustness due to integration properties. The study provides methodological references and empirical evidence for precision medicine prediction under limited labeling conditions.