Linguo-Statistical Analysis of Domain Concept Verbalization in the Russian Language
摘要
In this chapter, we examine the structure of language-independent domain knowledge in correlation with its verbalization in the domain lexicon and corpus of a particular language. The lexicon and corpus can be considered as domain models of language and speech in their dichotomy, thus permitting a comprehensive study of domain linguistic diversity and restrictions. The work is devoted to the specifics of the conceptual (categorical) division of the world in the “Research in athletes’ physiology” domain and verbalizations of the domain concepts in the Russian lexicon and corpus. The domain language-independent knowledge is represented by a multi-lingual ontology, while language-dependent knowledge is conveyed by a language-specific (Russian, in our case) lexicon, whose units are linked to the ontology concepts. Both the ontological and Russian domain lexical data are represented in the digital format and used as the knowledge component of a computer annotation tool. The latter allowed automating certain stages of the study and computing statistical indices of the analysis parameters. The work makes certain contribution to the development of the research methodology and computer instrument design. The novelty of the study also lies in the particular linguo-statistical analysis result values that can be used both for theoretical linguistic research and in applied aspects of natural language processing.