<p>The incorporation of real-time data from diverse sources, including wearables and sensors, into clinical decision-making presents a significant challenge. This challenge is primarily due to the limitations of traditional systems in processing high-velocity data streams and generating rapid, accurate predictions. This study proposes a scalable, distributed framework for real-time health monitoring and predictive analytics to address this gap. The system employs Apache Kafka for dependable data collection and Apache Spark for effective stream processing, facilitating real-time inference and offline model retraining. This study evaluates the framework utilizing two benchmark healthcare datasets: heart disease and diabetes. In the offline phase, historical data is employed to train and optimize various machine learning models, including Gradient Boosting, Random Forest, Decision Tree, Support Vector Machine, Logistic Regression, and Naive Bayes. The online phase utilizes pre-trained models to produce rapid predictions on real data streams, facilitating immediate responses in critical situations. The obtained results indicate that the Gradient Boosting model attained an accuracy of 98.10% in the offline phase and achieved 99.40% and 96.42% for heart disease and diabetes, respectively, in the online phase. Comparative analysis indicates that the proposed approach surpasses existing methods in terms of prediction accuracy and computational efficiency. This study introduces a comprehensive framework for the implementation of scalable, real-time predictive methods designed to improve the patients’ health monitoring.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Intelligent Healthcare: A Big Data and Machine Learning Approach for Real-Time Predictive Analytics

  • Mohammed Badawy,
  • Nagy Ramadan,
  • Hesham Ahmed Hefny

摘要

The incorporation of real-time data from diverse sources, including wearables and sensors, into clinical decision-making presents a significant challenge. This challenge is primarily due to the limitations of traditional systems in processing high-velocity data streams and generating rapid, accurate predictions. This study proposes a scalable, distributed framework for real-time health monitoring and predictive analytics to address this gap. The system employs Apache Kafka for dependable data collection and Apache Spark for effective stream processing, facilitating real-time inference and offline model retraining. This study evaluates the framework utilizing two benchmark healthcare datasets: heart disease and diabetes. In the offline phase, historical data is employed to train and optimize various machine learning models, including Gradient Boosting, Random Forest, Decision Tree, Support Vector Machine, Logistic Regression, and Naive Bayes. The online phase utilizes pre-trained models to produce rapid predictions on real data streams, facilitating immediate responses in critical situations. The obtained results indicate that the Gradient Boosting model attained an accuracy of 98.10% in the offline phase and achieved 99.40% and 96.42% for heart disease and diabetes, respectively, in the online phase. Comparative analysis indicates that the proposed approach surpasses existing methods in terms of prediction accuracy and computational efficiency. This study introduces a comprehensive framework for the implementation of scalable, real-time predictive methods designed to improve the patients’ health monitoring.