<p>Pattern recognition is the core discipline in artificial intelligence (AI) and machine learning (ML). It involves the identification and interpretation of patterns or regularities in data. This field has extensive applications across various domains, including natural language processing (NLP), image, and speech recognition. In natural language processing (NLP), pattern recognition techniques are pivotal for tasks such as text summarization, topic modeling, and sentiment analysis. Despite significant progress for languages with abundant resources, low-resource languages remain significantly under-represented, posing substantial challenges for NLP system development and deployment. This article introduces a brand-new sequence-to-sequence encoder-decoder model with Bahdanau attention that works especially well for summarizing complex ideas in Urdu. A new custom preprocessing system is introduced to handle script variations, morphological complexity, and diacritics, which makes the representation of text better. Unlike other methods, our model uses FastText embeddings to improve context understanding and a combination of extractive and abstractive summaries. A large Urdu dataset with one million articles covering a wide range of topics, such as business, science, entertainment, sports, and economics, was used to test our approach. We use ROUGE, BLEU, and semantic similarity metrics to test how well the suggested model works. The results show that attention processes work well in low-resource NLP applications, showing an improvement of up to 87% over baseline models. This research sets a new standard for summarizing in Urdu and shows how attention-based methods can be used with other languages that don’t have a lot of resources.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pattern matters: A deep learning approach with attention mechanism for text abstraction in low-ranked languages

  • Adil Ahmad,
  • Anwar Shah,
  • Bahar Ali,
  • Qamar Uz Zaman

摘要

Pattern recognition is the core discipline in artificial intelligence (AI) and machine learning (ML). It involves the identification and interpretation of patterns or regularities in data. This field has extensive applications across various domains, including natural language processing (NLP), image, and speech recognition. In natural language processing (NLP), pattern recognition techniques are pivotal for tasks such as text summarization, topic modeling, and sentiment analysis. Despite significant progress for languages with abundant resources, low-resource languages remain significantly under-represented, posing substantial challenges for NLP system development and deployment. This article introduces a brand-new sequence-to-sequence encoder-decoder model with Bahdanau attention that works especially well for summarizing complex ideas in Urdu. A new custom preprocessing system is introduced to handle script variations, morphological complexity, and diacritics, which makes the representation of text better. Unlike other methods, our model uses FastText embeddings to improve context understanding and a combination of extractive and abstractive summaries. A large Urdu dataset with one million articles covering a wide range of topics, such as business, science, entertainment, sports, and economics, was used to test our approach. We use ROUGE, BLEU, and semantic similarity metrics to test how well the suggested model works. The results show that attention processes work well in low-resource NLP applications, showing an improvement of up to 87% over baseline models. This research sets a new standard for summarizing in Urdu and shows how attention-based methods can be used with other languages that don’t have a lot of resources.