Abstract <p>The selection of papers related to a given topic involves forming an optimal order of the user working with primary sources from more general to more specific. This study is devoted to problems of ranking texts in this way on the basis of an analysis of their mutual semantic affinity by applying neural network language models of the Transformer architecture. Using these models, we map sentences of the text into multi-dimensional vectors called embeddings. Here, the estimation of semantic affinity (i.e., the “strength” of semantic relationship) of analyzed text fragments can be defined by the measure of closeness of vectors corresponding to them. The analyzed fragments of publications here are their abstracts along with titles. This paper resolves at the same time the relevant problem of complete description of the core content of the paper in its abstract and title. The offered solution uses the estimation developed by the authors for the text semantic coherence based on the closeness of vectors for its individual sentences. Herewith, the text of the abstract is extended subsequently by sentences of the introduction and conclusion of the article monitoring the variation in semantic coherence of extending abstracts according to the proposed estimation. To improve the accuracy of mutual estimation of the closeness of analyzed texts to the most rational (i.e., standard) variant of meaning transfer, an analogue for estimating the semantic coherence of an individual text with respect to text collection is introduced. The semantic coherence of the whole collection is defined by analogy with the estimation of the strength of the semantic relationship of the analyzed paper with the rest of the collection, and indicates how texts within the collection are mutually related in meaning. The final estimation of obtained results is carried out by ranking the collection according to closeness to the semantic standard using estimations proposed early by the authors in combination with clustering by applying the <i>k-</i>means method for embeddings of full texts of abstracts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Neural Network Language Modeling and Identification of Significant Fragments in Scientific Articles to Nonredundantly Transfer Their Meaning

  • D. V. Mikhaylov,
  • G. M. Emelyanov

摘要

Abstract

The selection of papers related to a given topic involves forming an optimal order of the user working with primary sources from more general to more specific. This study is devoted to problems of ranking texts in this way on the basis of an analysis of their mutual semantic affinity by applying neural network language models of the Transformer architecture. Using these models, we map sentences of the text into multi-dimensional vectors called embeddings. Here, the estimation of semantic affinity (i.e., the “strength” of semantic relationship) of analyzed text fragments can be defined by the measure of closeness of vectors corresponding to them. The analyzed fragments of publications here are their abstracts along with titles. This paper resolves at the same time the relevant problem of complete description of the core content of the paper in its abstract and title. The offered solution uses the estimation developed by the authors for the text semantic coherence based on the closeness of vectors for its individual sentences. Herewith, the text of the abstract is extended subsequently by sentences of the introduction and conclusion of the article monitoring the variation in semantic coherence of extending abstracts according to the proposed estimation. To improve the accuracy of mutual estimation of the closeness of analyzed texts to the most rational (i.e., standard) variant of meaning transfer, an analogue for estimating the semantic coherence of an individual text with respect to text collection is introduced. The semantic coherence of the whole collection is defined by analogy with the estimation of the strength of the semantic relationship of the analyzed paper with the rest of the collection, and indicates how texts within the collection are mutually related in meaning. The final estimation of obtained results is carried out by ranking the collection according to closeness to the semantic standard using estimations proposed early by the authors in combination with clustering by applying the k-means method for embeddings of full texts of abstracts.