A Study of Identification of Chinese VO Idioms with Statistical Measures
摘要
This paper describes a study of unsupervised identification of Chinese VO idioms by examining the Verb-Object (VO) pairs derived from the dependency structure of sentences. We test several statistical measures, including Point-wise Mutual Information (PMI), P(o|v), P(v|o), Salience, and Selectional Association. The experiments show that PMI performs the best in automatically identifying real VO idioms, which is consistent with previous studies on other languages. On the other hand, PMI tends to rank low-frequency items (very often noise) high. It obtained a 36% F1 score in the successful identification of real VO idioms among the top 100 of the ranked VO pairs. We thus suggest that syntactic features are not enough to identify VO idioms in an unsupervised framework, and more sophisticated methods with consideration of more semantic information are required.