Comparative Analysis of Pretrained Models for Text Classification, Generation and Summarization: A Detailed Analysis
摘要
The exponential growth in natural language processing (NLP) technologies has been propelled by the emergence of pretrained models, which have demonstrated remarkable efficacy across a spectrum of tasks including text classification, generation, and summarization. Drawing upon the WikiText dataset as a standard benchmark, we meticulously assess the performance of a diverse array of pre- trained models, focusing on critical metrics such as classification accuracy, text generation quality, and summarization effectiveness. Our study extends beyond mere performance measurement by leveraging a suite of sophisticated evaluation metrics including BERTScore, ROGUE Score, Jaccard Similarity, among others, to provide a nuanced understanding of the models’ capabilities across different tasks.Additionally, we employ the Technique for Order Preference by Similarity to an Ideal Solution (TOPSIS) method to aggregate the disparate performance metrics into a unified ranking framework, facilitating a comprehensive compar- ison of the pretrained models. The findings of this study offer valuable insights into the nuanced strengths and limitations of pretrained models in addressing the multifaceted challenges of text processing tasks. Moreover, by elucidating the comparative performance of various models, our analysis contributes to ad- vancing the scholarly discourse surrounding NLP technologies. For our Wikitext Dataset, GPT-3.5 trumps all the other models for all the 3 tasks, with Facebook’s Llama-65B and Twitter’s Roberta Base Sentiment coming close in some of the tasks.