Automatic generation of coherent and readable text in human language appears to be a promising technology in both academia and industry. Impressive results in this direction are achieved by large-scale pretrained language models capable of creating texts that are difficult to distinguish from human ones. However, controlling the generation process is a challenging task. Controllable text generation focuses on integrating user-specified constraints into the text generation process, such as a desired topic, style, or storyline. Approaches to control at the decoding stage of the language model have shown high efficiency for language control in comparison with resource-intensive fine-tuning of the model. In this work, we propose and investigate the COSMOS (COntrollable Semantic Method fOr Storytelling) method for generating a story from a storyline based on the semantic similarity of the decoded token sequences and the plot phrase. The idea of the method is to search for a sequence of language model tokens that is most semantically close to the required plot phrase. First, we generate several arbitrary small sequences of tokens from the autoregressive language model vocabulary. Then, using an autoencoding language model, we evaluate the semantic similarity of these sequences to the plot phrase and choose the sequence with the highest semantic similarity score. Experiments on the Russian-language corpus of fairy tales have shown the high efficiency of the proposed method for creating coherent stories that are highly relevant to a given storyline.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From Tokens to Tales: Semantic Similarity in Story Generation

  • Sergey Vychegzhanin,
  • Anastasia Kotelnikova,
  • Alexander Sergeev,
  • Evgeny Kotelnikov

摘要

Automatic generation of coherent and readable text in human language appears to be a promising technology in both academia and industry. Impressive results in this direction are achieved by large-scale pretrained language models capable of creating texts that are difficult to distinguish from human ones. However, controlling the generation process is a challenging task. Controllable text generation focuses on integrating user-specified constraints into the text generation process, such as a desired topic, style, or storyline. Approaches to control at the decoding stage of the language model have shown high efficiency for language control in comparison with resource-intensive fine-tuning of the model. In this work, we propose and investigate the COSMOS (COntrollable Semantic Method fOr Storytelling) method for generating a story from a storyline based on the semantic similarity of the decoded token sequences and the plot phrase. The idea of the method is to search for a sequence of language model tokens that is most semantically close to the required plot phrase. First, we generate several arbitrary small sequences of tokens from the autoregressive language model vocabulary. Then, using an autoencoding language model, we evaluate the semantic similarity of these sequences to the plot phrase and choose the sequence with the highest semantic similarity score. Experiments on the Russian-language corpus of fairy tales have shown the high efficiency of the proposed method for creating coherent stories that are highly relevant to a given storyline.