Extracting the Semantic Representation of Chinese-Japanese Homophones with Word2Vec for Teaching Chinese as a Second/Foreign Language
摘要
The purpose of this study was to quantify the semantic distance between Chinese-Japanese homophones, and to further examine whether the traditional classification of Chinese-Japanese homophones reflects varying semantic distances. Ultimately, this study aimed to offer corresponding teaching suggestions for teachers’ reference. We introduced a novel approach to classifying Chinese-Japanese homophones by combining the Revised Hierarchical Model with the Word2Vec technology to extract semantic representations and calculate the semantic distance of meanings in Chinese-Japanese homophones across “Same”, “Overlap”, and “Different” types. Through cluster analysis of semantic distances, we found that semantic distance indeed reflects the traditional classification of Chinese-Japanese homophones. The “Same” type had the average semantic distance of 0, showing the high degree of similarity between the meanings in Chinese and Japanese within this category. The “Different” type had the greatest average semantic distance, indicating the lowest similarity. The “Overlap” type had the average semantic distance between the “Same” and “Different” types, representing the complex nature of this category. Based on these findings, this study proposed teaching suggestions to help Chinese teachers determine the teaching sequence among Chinese-Japanese homophones.