Efficient Vietnamese Name Detection Using Highly Discriminative N-Grams
摘要
Detecting Vietnamese names from a global name collection is essential for connecting human resources within the Vietnamese community worldwide and for developing Vietnamese Expert Finding Systems. However, applying standard classification methods to recognize Vietnamese names is challenging due to their ambiguity with names from other countries. In this study, we proposed our novel approach, Highly Discriminative N-grams (HDN), which utilizes statistics from n-grams from human names for the classification task. We show that HDN is a lightweight and simple-to-reproduce approach for this task by benchmarking with other machine learning models.