<p>Using the 2021 Egypt Family Health Survey, this paper developed a logistic regression classifier, to predict children at risk of engaging in labor. Recognizing the inherent class imbalance within the child labor dataset, a comprehensive comparative analysis was undertaken to assess the effectiveness of multiple resampling techniques. The initial phase included forty-five experiments, comprising a baseline model (without resampling), twelve undersampling methods, eleven oversampling methods, and twenty-one filtering-based oversampling techniques. Subsequently, the top-performing techniques underwent further optimization by testing multiple parameter combinations, ending with an additional 180 experiments. The findings provide valuable insights into the profiles of children most vulnerable to engage in labor, contributing to a deeper understanding of this complex persistent issue. The key factors contributing to child labor, as identified by the classifier model, include children’s age group, geographical region of residence, poverty within families, mothers’ employment status, family land ownership, low levels of maternal education or lack thereof, and children not attending school. This predictive model holds potential as a practical tool for policymakers and researchers to design and implement targeted policy interventions effectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging machine learning and resampling techniques to analyze contributing factors to child labor in Egypt

  • Nahed T. Zeini,
  • Pakinam Mahmoud Fikry

摘要

Using the 2021 Egypt Family Health Survey, this paper developed a logistic regression classifier, to predict children at risk of engaging in labor. Recognizing the inherent class imbalance within the child labor dataset, a comprehensive comparative analysis was undertaken to assess the effectiveness of multiple resampling techniques. The initial phase included forty-five experiments, comprising a baseline model (without resampling), twelve undersampling methods, eleven oversampling methods, and twenty-one filtering-based oversampling techniques. Subsequently, the top-performing techniques underwent further optimization by testing multiple parameter combinations, ending with an additional 180 experiments. The findings provide valuable insights into the profiles of children most vulnerable to engage in labor, contributing to a deeper understanding of this complex persistent issue. The key factors contributing to child labor, as identified by the classifier model, include children’s age group, geographical region of residence, poverty within families, mothers’ employment status, family land ownership, low levels of maternal education or lack thereof, and children not attending school. This predictive model holds potential as a practical tool for policymakers and researchers to design and implement targeted policy interventions effectively.