错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detection of Smokers and Non-smokers Using Machine Learning

  • Nasreena Ali,
  • Yash Paul,
  • Amjad Husain,
  • Hemah Hussain

摘要

Smoking behavior is a complex and widespread health concern with significant implications for public health. According to research, smoking causes an estimated 8 million premature deaths each year. It is believed that 100 million individuals, mostly in wealthy nations, died prematurely throughout the twentieth century as a result of smoking. There are numerous ways in which smoking reduces productivity at work. Absenteeism and presenteeism are two immediate effects. In comparison to non-smokers, smokers typically take 31% more sick days and approximately three more sick days per year. The primary objective of this study is to classify smokers and non-smokers, to identify and explore the potential metrics required for the accurate detection of smokers and non-smokers, and to gain insights into the relationships between health factors and smoking patterns. The study employs a comprehensive dataset from Kaggle, containing features like demographic information and various health metrics, including blood pressure, cholesterol levels, and hemoglobin levels. Different machine learning algorithms, such as Random Forest, Extra Tree, Support Vector Machine, Logistic Regression, Gaussian Naive Bayes, Decision Tree, and K-Nearest Neighbors are applied to classify smokers and non-smokers based on the transformed health metrics. The models’ performances are evaluated using various metrics, including accuracy, precision, sensitivity, ROC-AUC, and confusion matrices. Our findings indicate that the Extra Tree classifier surpasses the performance of the other algorithms for one dataset, achieving Accuracy, Precision, sensitivity, specificity, and ROC-AUC scores of 87%, 84%, 92%, 82%, and 0.87, respectively. And for the second dataset, Logistic Regression outperformed the other algorithms with accuracy, precision sensitivity, specificity, and AUC of 97%, 86%, 100%, 96%, and 0.99, respectively. The implications of the study may be extended to personalized healthcare strategies and public health initiatives aimed at reducing smoking prevalence and its associated health risks.