错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Summarization of Telugu Text Discourses

  • K. Bhuvaneshwari,
  • S. A. JyothiRani,
  • V. V. Haragopal

摘要

A text summary is a condensed version of the original text that highlights the key points. Most of the Natural Language processing applications are widely used on text data which is readily available in English language. Limited research is being carried out on Telugu language textual data. Manual text summarization necessitates a large group of talented, unbiased persons, a substantial financial investment, and a considerable amount of time. In this study, we proposed an investigative approach for extractive text summarization that uses essential features such as order of the sentences in the given document, sentence similarities with title, word-frequencies, and document centrality of Telugu language to summarize the text data in Telugu. To rank the sentences, an improved sentence scoring approach is used. The event and named entity scores are used in this sentence scoring method. Applying statistical measures to the events and named entities that have been obtained is how sentence scoring is done. The word frequency wf score is determined by counting the instances of the event or named entity. The inverse sentence frequency isf can be calculated using the number of sentences in which the events and named entities occurred. The wf isf score of term t is calculated as the product of word frequency and inverse sentence frequency. Finally, based on the wf isf score, it determines the terms significance and includes it in the summary. The sentences were then ranked based on the scores calculated for each and every phrase while taking into account all of the given features. For this research study, we collected Telugu speech transcripts of a proficient Indian speaker Chaganti Koteswara Rao, recognized for his talks on Sanatana Dharma.