The research was mostly motivated by the extensive theoretical demand for understanding the notion of text complexity and text profiling, on the one hand, but limited research on complexity of Russian texts, on the other. The article presents RuLingva, a text profiler for the Russian language, and offers readers results of exploring its functions as a discriminator of informative and instructional texts. We offer a brief overview of similar tools for English (Coh-Metrics, TextInspector, TAACO, LIWC, TAALES), and Russian (Textometr) focusing on 49 parameters measured by RuLingva. The list includes descriptive (length in types, tokens, syllables), phonological (number of mono-, di-, three and four-syllable types and tokens), morphological (POS, lists of and numbers of content words, numbers of nominal and verbal categories manifestations), lexical (abstractness, frequency, TTR) and discourse (local and global noun and argument overlap) parameters. We also present results of an experimental test of RuLingva’s functions and demonstrate that it can successfully identify a text type provided it is supplied with ranges of reference parameters. We also outline the areas of RuLingva application demonstrating that the service can be installed into Russian search engines with the aim to stream text types thus narrowing the range or sources and simplifying information retrieval. RuLingva is in great demand among speechwriters, text analytics, teachers, test developers, as well as textbooks authors, politicians and journalists.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multilevel Analyses of Russian Texts with RuLingva: A Case Study

  • Marina Solnyshkina,
  • Valery Solovyev,
  • Andrew Danilov,
  • Radif Zamaletdinov,
  • Svetlana Akhtyamova

摘要

The research was mostly motivated by the extensive theoretical demand for understanding the notion of text complexity and text profiling, on the one hand, but limited research on complexity of Russian texts, on the other. The article presents RuLingva, a text profiler for the Russian language, and offers readers results of exploring its functions as a discriminator of informative and instructional texts. We offer a brief overview of similar tools for English (Coh-Metrics, TextInspector, TAACO, LIWC, TAALES), and Russian (Textometr) focusing on 49 parameters measured by RuLingva. The list includes descriptive (length in types, tokens, syllables), phonological (number of mono-, di-, three and four-syllable types and tokens), morphological (POS, lists of and numbers of content words, numbers of nominal and verbal categories manifestations), lexical (abstractness, frequency, TTR) and discourse (local and global noun and argument overlap) parameters. We also present results of an experimental test of RuLingva’s functions and demonstrate that it can successfully identify a text type provided it is supplied with ranges of reference parameters. We also outline the areas of RuLingva application demonstrating that the service can be installed into Russian search engines with the aim to stream text types thus narrowing the range or sources and simplifying information retrieval. RuLingva is in great demand among speechwriters, text analytics, teachers, test developers, as well as textbooks authors, politicians and journalists.