Speech Technology is increasingly popular due to the high demands of voice-activated systems and security purposes. As the volume of speech data grows, obtaining important information becomes increasingly difficult. There comes the need to generate information-rich summaries. Even with high-effort research, the results have become stagnant over the years. The current study aims to deeply analyze the shortcomings in the available speech summarization techniques. It was found that this could be due to the homogeneous datasets, unconstrained reference summaries, structural bias in the training data, limitations of the evaluation metrics, or the papers’ limited discussion of implementation specifics. After analyzing the findings from several papers, it was found that these drawbacks could be addressed by 1) limiting the “inverted pyramid” structure in the training data, 2) introducing new bias-free datasets with constrained summaries across domains and languages, 3) implementing evaluation metrics that are in line with user preferences, or 4) closely examining implementation details. This would immensely help in choosing, building, or modifying speech summarization tools in the future.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speech Summarization Diverse Challenges and Advancements

  • Riya Ahlawat,
  • Samriddhi Tiwari,
  • Muskan Singh,
  • Shweta Jindal

摘要

Speech Technology is increasingly popular due to the high demands of voice-activated systems and security purposes. As the volume of speech data grows, obtaining important information becomes increasingly difficult. There comes the need to generate information-rich summaries. Even with high-effort research, the results have become stagnant over the years. The current study aims to deeply analyze the shortcomings in the available speech summarization techniques. It was found that this could be due to the homogeneous datasets, unconstrained reference summaries, structural bias in the training data, limitations of the evaluation metrics, or the papers’ limited discussion of implementation specifics. After analyzing the findings from several papers, it was found that these drawbacks could be addressed by 1) limiting the “inverted pyramid” structure in the training data, 2) introducing new bias-free datasets with constrained summaries across domains and languages, 3) implementing evaluation metrics that are in line with user preferences, or 4) closely examining implementation details. This would immensely help in choosing, building, or modifying speech summarization tools in the future.