Overview of Approaches for Increasing Coherence in Extractive Summaries
摘要
This paper provides a comprehensive overview of various methodologies aimed at enhancing coherence in summaries generated through extractive summarization techniques. With the exponential growth of information, particularly in the scientific domain, the need for effective summarization tools is more pressing than ever. We explore several approaches including graph-based, knowledge-based, Latent Semantic Analysis (LSA), Hidden Markov Model (HMM), and discourse structure-based methods. We also consider the role of datasets like the Grammarly Corpus of Discourse Coherence (GCDC) and a dataset of extractive and corresponding abstractive summaries of scientific papers. Our analysis of these methodologies and datasets aims to shed light on the current state of the field and contribute to the ongoing efforts to improve the coherence of automatic summaries. The paper concludes with the assertion that despite the challenges, extractive summarization methods still hold significant potential for further development.