Abstract <p>The paper proposes a solution for generating image captions in both formal and conversational registers. This study is motivated by the need for educational tools to assist non-native speakers in mastering colloquial Russian. The methodology employs a multimodal encoder–decoder ensemble architecture, in which a pre-trained ResNet-152 Convolutional Neural Network serves as the encoder, and an LSTM network functions as the decoder. The captioning performance is further enhanced by incorporating the Bahdanau attention mechanism. To facilitate training, the authors constructed a proprietary dataset derived from MS COCO, which was translated and stylistically adapted via the GigaChat large language model. During ensemble construction, ruCLIPScore is utilized to select the most effective model configurations. Experimental results indicate that the ensemble significantly outperforms its individual constituent models according to ruCLIPScore and can produce captions with stylistic diversity across registers.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generating Russian-Language Image Descriptions in Both Formal and Conversational Styles via a Neural-Network Ensemble

  • M. A. Privalov,
  • A. S. Kozharinov

摘要

Abstract

The paper proposes a solution for generating image captions in both formal and conversational registers. This study is motivated by the need for educational tools to assist non-native speakers in mastering colloquial Russian. The methodology employs a multimodal encoder–decoder ensemble architecture, in which a pre-trained ResNet-152 Convolutional Neural Network serves as the encoder, and an LSTM network functions as the decoder. The captioning performance is further enhanced by incorporating the Bahdanau attention mechanism. To facilitate training, the authors constructed a proprietary dataset derived from MS COCO, which was translated and stylistically adapted via the GigaChat large language model. During ensemble construction, ruCLIPScore is utilized to select the most effective model configurations. Experimental results indicate that the ensemble significantly outperforms its individual constituent models according to ruCLIPScore and can produce captions with stylistic diversity across registers.