Natural Language Processing, a field under Data Science, has helped us to retrieve a quick summary of large texts like news articles, stories, etc., so that we do not have to spend a lot of time in reading long paragraphs. Over the years, abstractive text summarization has been done using various Deep Learning models, but we started achieving better results after the introduction of the transformers and self-attention mechanisms, as they were able to better understand the human language. In the recent times, the use of Large Language Models (LLMs) has surfaced for a variety of NLP tasks. They are pre-trained on much larger corpus and generate results very close to what the humans would generate. In this paper, we have performed a comparative analysis of the working, evaluation, and the differences with the Pre-trained Encoder-Decoder Models and Large Language Models to generate Abstract Text Summaries and have analysed some key points about working with LLMs. Our paper works as the baseline reference for the transition from the Pre-trained Encoder-Decoder models to LLMs. We have also presented the results of the LLaMA2-13b-chat model on the CNN/DailyMail Corpus for Abstractive Text Summarization with the rouge-1 value of 37.9, which has shown a significant improvement than the previous research work done with LLMs like text-davinci-003 and GPT3-D2.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Study on the Working of Pre-trained Encoder-Decoder Models and Large Autoregressive Language Models for Abstractive Text Summarization

  • Niranjana Sowpari,
  • Poonam Bansal

摘要

Natural Language Processing, a field under Data Science, has helped us to retrieve a quick summary of large texts like news articles, stories, etc., so that we do not have to spend a lot of time in reading long paragraphs. Over the years, abstractive text summarization has been done using various Deep Learning models, but we started achieving better results after the introduction of the transformers and self-attention mechanisms, as they were able to better understand the human language. In the recent times, the use of Large Language Models (LLMs) has surfaced for a variety of NLP tasks. They are pre-trained on much larger corpus and generate results very close to what the humans would generate. In this paper, we have performed a comparative analysis of the working, evaluation, and the differences with the Pre-trained Encoder-Decoder Models and Large Language Models to generate Abstract Text Summaries and have analysed some key points about working with LLMs. Our paper works as the baseline reference for the transition from the Pre-trained Encoder-Decoder models to LLMs. We have also presented the results of the LLaMA2-13b-chat model on the CNN/DailyMail Corpus for Abstractive Text Summarization with the rouge-1 value of 37.9, which has shown a significant improvement than the previous research work done with LLMs like text-davinci-003 and GPT3-D2.