<p>Many scholars have traditionally examined features linked to students’ science performance by using classical statistical methods. However, there is a dearth of research on predicting student achievements by using machine learning (ML) models in science education. We aim to address this gap using the 2006 and 2015 Program for International Student Assessment datasets to construct prediction models for students’ scientific literacy. We applied three models (traditional regression, XGBoost, and random forest) to compare their accuracy. The two ML models showed more accuracy and interpretability than the regression model. For the overall group, the results of two ML models indicated that grade, enjoyment, and self-perceived scientific literacy were the most important predictors of scientific literacy. For both low and high performers, the results of the two ML models were similar to the overall group regardless of socioeconomic status. In both ML models, motivation was the least important variable overall, yet it became the most significant predictor in the regression model. Enjoyment, which exerted the highest influence in the two ML models, was the least significant predictor in the regression model. This study highlights the benefits of integrating ML models into predictive analyses in science education.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning Approaches for Predicting U.S. Students’ Scientific Literacy: An Analysis of Key Factors Across Performance Levels and Socioeconomic Statuses

  • Hyesun You,
  • Minju Hong,
  • Li Zhu,
  • Fang Zhenhan

摘要

Many scholars have traditionally examined features linked to students’ science performance by using classical statistical methods. However, there is a dearth of research on predicting student achievements by using machine learning (ML) models in science education. We aim to address this gap using the 2006 and 2015 Program for International Student Assessment datasets to construct prediction models for students’ scientific literacy. We applied three models (traditional regression, XGBoost, and random forest) to compare their accuracy. The two ML models showed more accuracy and interpretability than the regression model. For the overall group, the results of two ML models indicated that grade, enjoyment, and self-perceived scientific literacy were the most important predictors of scientific literacy. For both low and high performers, the results of the two ML models were similar to the overall group regardless of socioeconomic status. In both ML models, motivation was the least important variable overall, yet it became the most significant predictor in the regression model. Enjoyment, which exerted the highest influence in the two ML models, was the least significant predictor in the regression model. This study highlights the benefits of integrating ML models into predictive analyses in science education.