Predictive Modeling for Food Security Assessment Using Synthetic Minority Over-Sampling Technique
摘要
In Pakistan food insecurity remains a major public health issue. This work develops and test predictive models using machine learning techniques for household checking levels of Food Security (FS). In addition, the paper will analyze food consumption scores and WASH indicators to determine an innovative method for predicting households’ level of food security in various regions across Pakistan. This work uses an integrated approach to comprehensively investigate various factors affecting food insecurity. A two-stage cluster sampling was used, and data collected by a mobile tool in standardized questionnaire. Performance metrics for predicting FS using various machine learning models are evaluated. This work also describes the strengths and limitations of each model. Notably, the Random Forest model achieved an impressive accuracy of 99.86%, demonstrating its superior ability to handle the complexities of food security data. Logistic Regression performs well on this data and the performance of our model indicates that it is doing reasonably good at cross-validation (stable validation results), suggesting confidence in generalizing new samples. The research will help to define a ground for accommodating WASH data in FS assessment, which might be useful for policymaking and intervention strategies oriented towards health- and nutrition-related domains.