错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Relevance of Sentence Features for Multi-document Text Summarization Using Human-Written Reference Summaries

  • Verónica Neri Mendoza,
  • Yulia Ledeneva,
  • René Arnulfo García-Hernández,
  • Ángel Hernández Castañeda

摘要

For multi-document text summarization, text features are fundamental because they determine the importance of each sentence from source documents. Therefore, selected sentences create a summary that represents the most essential information. In the state-of-the-art, several techniques and methods have been proposed that use different text features and select sentences. However, some features may be more important than others. Thus, differentiating between important and unimportant features is a difficult task. This work proposes a method to generate extractive multi-document text summaries based on statistical and linguistic text features. We calculated the relevance coefficient of each feature to determine its degree of importance through the human-written reference summaries. To perform such calculus, we use 19 text features. After this calculus, we employ a Genetic Algorithm (GA) that selects sentences to generate summaries. In a general way, the proposed method consists of three steps: feature weighting, concatenation and preprocessing of source documents, and feature extraction with sentence selection. In our experiments, we used the DUC01 dataset in two different lengths to evaluate the performance of the proposed method. The results show improvement over state-of-the-art methods.