Empowering Multilingual Abstractive Text Summarization: A Comparative Study of Word Embedding Techniques
摘要
Multilingual Abstractive Text Summarization is one of the many critical Natural Language Processing tasks that involves generating concise and coherent summaries in multiple languages. The use of word embeddings has emerged as a pivotal technique in this domain, as it enables the representation of words as continuous vectors, capturing semantic relationships and contextual information. This paper presents an in-depth exploration of the role of word embeddings in enhancing Multilingual Abstractive Text Summarization, specifically focusing on Hindi and Marathi languages. This study investigates various word embedding techniques, including ELMo, BERT, and XLNet, and their impact on the quality of abstractive summarization for Hindi and Marathi texts by conducting comprehensive experiments using deep learning-based summarization models trained on diverse datasets in the target languages. Through rigorous evaluation, the performance of each word embedding technique using metrics such as ROUGE (Recall-Oriented Understudy for Gisting Evaluation) was analyzed. Findings of this paper reveal that the choice of word embeddings significantly influences the summarization quality for both Hindi and Marathi languages. Certain embeddings excel in capturing the linguistic nuances and semantic representations specific to each language, resulting in more coherent and informative summaries. Furthermore, the efficacy of multilingual word embeddings in cross-lingual summarization is explored, demonstrating promising results in preserving semantic relationships across Hindi and Marathi, leading to improved summarization outcomes.