A Hybrid Framework for Implementing Modified K-Means Clustering Algorithm for Hindi Word Sense Disambiguation
摘要
In this paper, we have proposed a new method for the implementation of an unsupervised K-Means clustering algorithm for Word Sense Disambiguation in Hindi language. We created a feature vector for each instance of an ambiguous target word, as well as for the lexical information obtained by utilizing Hindi WordNet for the same word. The lexical information was obtained from synset definitions, and its glosses for a given ambiguous target word from Hindi WordNet. The seed value was selected randomly from the lexical information. The features used for the experiment were term-frequency, term-frequency inverse instance-frequency, and n-grams. We applied similarity measures such as Cosine, Jaccard, and Dice similarity measures. We performed an experiment on the Hindi dataset containing three polysemous nouns. We obtained an average accuracy of 69.6%. We also observed that the accuracy of Dice similarity measure performs better with the modified K-means clustering algorithm.