Efficient Text Analysis: A BERT-Based Approach to Named Entity Recognition (NER) and Classification for Malayalam Language
摘要
The onset of the present millennium has witnessed a remarkable surge in the use of Indian language in various forms of multilingual interpersonal communication, including news forums, webpages, email, and social media conversations. The extensive dissemination of mobile phones has ignited an apparently unceasing utilization of well-known micro-blogging platforms—Facebook, Twitter, and Watsapp—to express thoughts, remarks, and impressions in the field of natural language processing (NLP). The primary objective of this research is to investigate Efficient Text Analysis using Bidirectional Encoder Representations from Transformers (BERT) Based for Named Entity Recognition and Classification in the Malayalam language. To conduct Named Entity Recognition (NER) in the Malayalam dialect, which is widely spoken in Kerala, India, this study employs both deep learning and the BERT approach. It involves customization of a pre-trained multilingual BERT (m-BERT) model to accurately capture the unique phonetic features of Malayalam. Lastly, the performance of the proposed approach is validated and compared with existing recent studies based on evaluation metrics such as accuracy (%), precision (%), recall (%), and f1-score (%). The results show that the proposed approach attains the highest accuracy (97.10%), thereby outperforming the existing recent studies.