<p>People often fail to notice chronic kidney disease (CKD) when it is in the early stages, risking entrance into late stages that can be difficult to cure. To mitigate this problem, we hereby propose an online machine learning (ML) system to help predict CKD early enough using the omnipresent social media platform <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="43994_2025_269_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="13" /> </InlineMediaObject> <EquationSource Format="TEX">\(\mathbb {X}\)</EquationSource> </InlineEquation>. Through the system, a person can submit a post containing a predefined hashtag and a record of medical features, and get in return an immediate positive/negative diagnosis. The post is retrieved in the system by a Kafka topic, and is then sent to Spark Streaming, where it is encapsulated as a feature vector ready to enter an ML prediction model. The model is chosen from several models trained and tested on CKD datasets. To accommodate huge post streams, Apache Spark is employed as a computational engine. The system has been implemented in Python and PySpark for validation and performance evaluation, where performance is measured in terms of accuracy, precision and F-Score. For credible results, 10-fold cross validation is employed. To ensure efficient and accurate prediction, feature selection is carried out using <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="43994_2025_269_Article_IEq3.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\chi ^2\)</EquationSource> </InlineEquation>, <i>F</i>-statistic and mutual information (MI). To enhance performance further, hyperparameter tuning is used. The experimental results show that the best model to predict CKD within the described system is random forest (RF), which exhibited when tested on the Kaggle dataset an impressive accuracy of 99.14%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Design and performance analysis of an intelligent system for predicting chronic kidney disease (CKD) in real time using the \(\mathbb {X}\) platform

  • Tarneem Elemam,
  • Mohamed Abdelsabour Fahmy,
  • Hamed Nassar

摘要

People often fail to notice chronic kidney disease (CKD) when it is in the early stages, risking entrance into late stages that can be difficult to cure. To mitigate this problem, we hereby propose an online machine learning (ML) system to help predict CKD early enough using the omnipresent social media platform \(\mathbb {X}\) . Through the system, a person can submit a post containing a predefined hashtag and a record of medical features, and get in return an immediate positive/negative diagnosis. The post is retrieved in the system by a Kafka topic, and is then sent to Spark Streaming, where it is encapsulated as a feature vector ready to enter an ML prediction model. The model is chosen from several models trained and tested on CKD datasets. To accommodate huge post streams, Apache Spark is employed as a computational engine. The system has been implemented in Python and PySpark for validation and performance evaluation, where performance is measured in terms of accuracy, precision and F-Score. For credible results, 10-fold cross validation is employed. To ensure efficient and accurate prediction, feature selection is carried out using \(\chi ^2\) , F-statistic and mutual information (MI). To enhance performance further, hyperparameter tuning is used. The experimental results show that the best model to predict CKD within the described system is random forest (RF), which exhibited when tested on the Kaggle dataset an impressive accuracy of 99.14%.