The present study introduces a health insurance prediction system that leverages machine learning methodologies. In contemporary times, there has been a notable increase in endeavors focused on tackling this matter, since the significance of health insurance as a research topic has markedly escalated following the occurrence of the pandemic. The dataset employed in this research comprises 1338 observations, 7 columns, and corresponds to individual medical expenditures in the United States which is available at the Kaggle platform. The dataset encompasses a variety of variables utilized in the prediction of insurance prices, including age, gender, BMI, smoking status, and number of children. The researchers utilized machine learning models including neural network, XAI, auto modeling to determine the correlation between pricing and the aforementioned attributes. The training process involved partitioning the dataset into a 80–20 ratio for training and evaluation respectively. As a consequence, the system achieved an accuracy rate of 97% by Gradient Boosting but we corrected it as 92% by Gradient Boosting Regressor by encoding and hyper tuning. Also, among predictive machine learning models, Random Forest had the best accuracy i.e., of 83.44%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Healthcare Cost Patterns and Prediction: Investigating Personal Datasets Using Data Analytics

  • Md. Aminul Islam,
  • Anindya Nag,
  • Pretam Chandra,
  • Bhupesh Kumar Mishra,
  • S. M. Firoz Ahmed Fahim,
  • Md Mozammel Hoque

摘要

The present study introduces a health insurance prediction system that leverages machine learning methodologies. In contemporary times, there has been a notable increase in endeavors focused on tackling this matter, since the significance of health insurance as a research topic has markedly escalated following the occurrence of the pandemic. The dataset employed in this research comprises 1338 observations, 7 columns, and corresponds to individual medical expenditures in the United States which is available at the Kaggle platform. The dataset encompasses a variety of variables utilized in the prediction of insurance prices, including age, gender, BMI, smoking status, and number of children. The researchers utilized machine learning models including neural network, XAI, auto modeling to determine the correlation between pricing and the aforementioned attributes. The training process involved partitioning the dataset into a 80–20 ratio for training and evaluation respectively. As a consequence, the system achieved an accuracy rate of 97% by Gradient Boosting but we corrected it as 92% by Gradient Boosting Regressor by encoding and hyper tuning. Also, among predictive machine learning models, Random Forest had the best accuracy i.e., of 83.44%.