Abstract <p>The article discusses the use of stylometric<sup>1</sup> analysis in plagiarism detection in texts in the Tatar language. Relevant tools have been developed, utilizing machine learning algorithms, including clustering (<i>k</i>‑means clustering), classification (random forest method, support vector machine method, naïve Bayes classifier), and a hybrid approach (FastText model + logistic regression). Special attention is paid to the adaptation of linguistic metrics to the Tatar language. The possibility is demonstrated of using stylometric analysis methods to address tasks of authorship attribution and determination of style and emotional tone in texts in the Tatar language.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Stylometric Analysis in the Task of Plagiarism Detection in Texts in the Tatar Language

  • I. Z. Khayaleeva,
  • M. M. Abramskiy

摘要

Abstract

The article discusses the use of stylometric1 analysis in plagiarism detection in texts in the Tatar language. Relevant tools have been developed, utilizing machine learning algorithms, including clustering (k‑means clustering), classification (random forest method, support vector machine method, naïve Bayes classifier), and a hybrid approach (FastText model + logistic regression). Special attention is paid to the adaptation of linguistic metrics to the Tatar language. The possibility is demonstrated of using stylometric analysis methods to address tasks of authorship attribution and determination of style and emotional tone in texts in the Tatar language.