Event Uncertainty for Twitter Data Using Thematic Context Vector
摘要
One of the critical steps in text mining is keyword extraction. Extracting keywords along with their respective contextual events in Twitter data poses a significant challenge due to language informality, encompassing acronyms, misspelled words, synonyms, transliteration, and ambiguous terms. The current systems for keyword extraction rely on either pattern-based or event-based approaches. In this paper, keywords are extracted using contextual events with the assistance of data association. The thematic context vectors for events are identified using the uncertainty principle in the proposed system. The system is tested on the Twitter COVID-19 dataset and proves to be effective. It extracts event-specific thematic context vectors from the test dataset and ranks them. The extracted thematic context vectors are then utilized for the clustering of contextual thematic vectors, improving the silhouette coefficient by 0.5% compared to state-of-the-art methods, namely TF and TF-IDF. The thematic context vector can be applied to other applications such as cyberbullying detection, sarcasm detection, and figurative language detection.