Artificial intelligence in linguistics: a GBRT model approach to forecast Cantonese levels among Chinese Malaysians
摘要
This study leverages a Gradient Boosted Regression Trees (GBRT) machine learning model to explore how Cantonese media exposure and cultural identity affect Cantonese language proficiency among Chinese Malaysians. By integrating sociolinguistic insights with predictive modeling, we address the multidimensional nature of language use factors. Using survey data from 642 Chinese Malaysian respondents, the GBRT model achieved a high predictive accuracy (R² ≈ 0.90) for Cantonese proficiency. The model identified key predictors, such as daily Cantonese use in social settings, media engagement, and generational cohort, underscoring their significant roles in language maintenance. These findings demonstrate the potential of machine learning to advance sociolinguistic research and provide practical insights for preserving linguistic heritage in multicultural societies.