Detecting Vietnamese names from a global name collection is essential for connecting human resources within the Vietnamese community worldwide and for developing Vietnamese Expert Finding Systems. However, applying standard classification methods to recognize Vietnamese names is challenging due to their ambiguity with names from other countries. In this study, we proposed our novel approach, Highly Discriminative N-grams (HDN), which utilizes statistics from n-grams from human names for the classification task. We show that HDN is a lightweight and simple-to-reproduce approach for this task by benchmarking with other machine learning models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Vietnamese Name Detection Using Highly Discriminative N-Grams

  • Toan Luu,
  • Phan Quoc Hung Mai,
  • Thi Thu Phuong Dao,
  • Quang Hung Nguyen,
  • Xuan Lam Pham

摘要

Detecting Vietnamese names from a global name collection is essential for connecting human resources within the Vietnamese community worldwide and for developing Vietnamese Expert Finding Systems. However, applying standard classification methods to recognize Vietnamese names is challenging due to their ambiguity with names from other countries. In this study, we proposed our novel approach, Highly Discriminative N-grams (HDN), which utilizes statistics from n-grams from human names for the classification task. We show that HDN is a lightweight and simple-to-reproduce approach for this task by benchmarking with other machine learning models.