Graphical abbreviation is a method of shortening words to save time and space on the page. This paper explores the distinction between graphical abbreviations and acronyms and provides a classification system for different types of abbreviations. We focus on the development and evaluation of a novel model designed for automatic abbreviation expansion in Russian—a task complicated by the language’s rich inflectional morphology. Our approach combines dictionary-based methods with a masked language model (BERT) to handle both unambiguous and ambiguous abbreviations while ensuring the correct expansion based on grammatical context. To further enhance model performance, we augmented a custom dataset with examples from news articles, providing diverse contexts for abbreviation use. We evaluate the model’s performance using perplexity, Word Error Rate (WER), and Lemma Error Rate (LER). Our results indicate that the model achieves high accuracy, with a WER of 0.0376 and an LER of 0.0299, making it highly effective for practical applications in natural language processing tasks such as text normalization, automatic speech recognition, and digital assistants. This research highlights the importance of context-aware models in abbreviation expansion, particularly for languages with complex morphological systems like Russian.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Graphical Abbreviation Disclosure in Russian Language

  • Stanislav Elkin

摘要

Graphical abbreviation is a method of shortening words to save time and space on the page. This paper explores the distinction between graphical abbreviations and acronyms and provides a classification system for different types of abbreviations. We focus on the development and evaluation of a novel model designed for automatic abbreviation expansion in Russian—a task complicated by the language’s rich inflectional morphology. Our approach combines dictionary-based methods with a masked language model (BERT) to handle both unambiguous and ambiguous abbreviations while ensuring the correct expansion based on grammatical context. To further enhance model performance, we augmented a custom dataset with examples from news articles, providing diverse contexts for abbreviation use. We evaluate the model’s performance using perplexity, Word Error Rate (WER), and Lemma Error Rate (LER). Our results indicate that the model achieves high accuracy, with a WER of 0.0376 and an LER of 0.0299, making it highly effective for practical applications in natural language processing tasks such as text normalization, automatic speech recognition, and digital assistants. This research highlights the importance of context-aware models in abbreviation expansion, particularly for languages with complex morphological systems like Russian.