Abstract <p>One of the key elements in solving spam filtering problems is the text vectorization method. This article proposes a vectorization method based on matching text to pairs of intentionalities. A list of intentionality pairs is extracted and a synthetic dataset is generated from text utterances. A neural network is designed and trained to determine the degree to which each intentionality belongs to the textual expression provided as the input of the model. The developed method is tested on the problem of filtering spam messages using logistic regression and the Enron dataset and SMS dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Construction of a Semantic Space of Intentionalities Using Generative Pretrained Models to Solve the Problem of Spam Filtering

  • I. Yu. Zhukov,
  • E. E. Balashova,
  • A. P. Mandrov,
  • N. D. Kravchenko

摘要

Abstract

One of the key elements in solving spam filtering problems is the text vectorization method. This article proposes a vectorization method based on matching text to pairs of intentionalities. A list of intentionality pairs is extracted and a synthetic dataset is generated from text utterances. A neural network is designed and trained to determine the degree to which each intentionality belongs to the textual expression provided as the input of the model. The developed method is tested on the problem of filtering spam messages using logistic regression and the Enron dataset and SMS dataset.