An Unsupervised Artificial Intelligence Strategy for Recognising Multi-word Expressions in Transformed Bengali Data
摘要
Multiword expressions are linguistic phrases that are both distinctive to each language yet universal. Handling multiword expressions is crucial in natural language processing applications since it presents several challenges. This article describes the process of identifying multiword phrases in the investigated Bengali corpus using association rule mining algorithms. The proposed system executes ten procedures in three phases: preprocessing, extraction of frequent item sets, accuracy evaluation, and validation. This experiment used a restricted amount of data since extra computer resources and standard multiword phrases were needed to confirm the computational findings. The ideas of entropy and Naive Bayes are used to validate the results of this experiment, which has undergone extensive investigation. This experiment was primarily carried out by employing a few phrases from the Agriculture domain of the experimented corpus. Based on validation using the Naive Bayes theorem, our experiment reveals that top-ranked terms have a high possibility of producing multiword expressions. Finally, we accomplished phrases with a high likelihood of becoming multiword expressions as a result of the n-gram combinations.