<p>This study leverages a Gradient Boosted Regression Trees (GBRT) machine learning model to explore how Cantonese media exposure and cultural identity affect Cantonese language proficiency among Chinese Malaysians. By integrating sociolinguistic insights with predictive modeling, we address the multidimensional nature of language use factors. Using survey data from 642 Chinese Malaysian respondents, the GBRT model achieved a high predictive accuracy (R² ≈ 0.90) for Cantonese proficiency. The model identified key predictors, such as daily Cantonese use in social settings, media engagement, and generational cohort, underscoring their significant roles in language maintenance. These findings demonstrate the potential of machine learning to advance sociolinguistic research and provide practical insights for preserving linguistic heritage in multicultural societies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Artificial intelligence in linguistics: a GBRT model approach to forecast Cantonese levels among Chinese Malaysians

  • Yuqing Peng,
  • Junxian Xie,
  • Lin Zhang,
  • Yuwen Lyu

摘要

This study leverages a Gradient Boosted Regression Trees (GBRT) machine learning model to explore how Cantonese media exposure and cultural identity affect Cantonese language proficiency among Chinese Malaysians. By integrating sociolinguistic insights with predictive modeling, we address the multidimensional nature of language use factors. Using survey data from 642 Chinese Malaysian respondents, the GBRT model achieved a high predictive accuracy (R² ≈ 0.90) for Cantonese proficiency. The model identified key predictors, such as daily Cantonese use in social settings, media engagement, and generational cohort, underscoring their significant roles in language maintenance. These findings demonstrate the potential of machine learning to advance sociolinguistic research and provide practical insights for preserving linguistic heritage in multicultural societies.