<p>Image captioning endeavors to generate text descriptions from images using the learned cross-modal generator. Automatically generating descriptions for life ecological experiment images of China’s space station can significantly expedite the comprehension of the intricate semantic details within the experimental data, facilitating the efficient management and accurate retrieval of massive experimental data. However, current image caption approaches trained on general image caption datasets cannot be directly applied to life ecological experiment scenes. In this work, we construct a Chinese caption dataset of life ecological experiment images by collecting and annotating the life ecological experiment image products of China’s space station (LEEoCSSCc). Furthermore, we propose an innovative image captioning approach by exploiting cross-modal prediction and data augmentation (IACPDA), which is designed to harness the raw image to guide the semantic constraints of the generated captions from different perspectives. IACPDA facilitates the transformation of both the original image and the resulting caption into a unified semantic space. It then leverages the predictions derived from the original image to offer constructive guidance for the generated captions. Moreover, we use a semantic-guided image augmentation method based on a diffusion model to generate augmented images that are distinct from the original training images while maintaining the essential semantics of the original images. Consequently, these augmented images can be effectively utilized as soft labels, enabling the learning of a more effective generator. The experiments indicate that our method achieves superior performance compared to other state-of-the-art baselines on the LEEoCSSCc. This paper is the first application of image captioning to the analysis of space science experiment images from China’s space station and further offers vital technical support for the understanding of space science experiments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image captioning for life ecological experiment of China’s space station

  • Yunfei Liu,
  • Yizhao Wang,
  • Chen Du,
  • Yunziwei Deng,
  • Anqi Liu,
  • Yanan Liu,
  • ShengYang Li

摘要

Image captioning endeavors to generate text descriptions from images using the learned cross-modal generator. Automatically generating descriptions for life ecological experiment images of China’s space station can significantly expedite the comprehension of the intricate semantic details within the experimental data, facilitating the efficient management and accurate retrieval of massive experimental data. However, current image caption approaches trained on general image caption datasets cannot be directly applied to life ecological experiment scenes. In this work, we construct a Chinese caption dataset of life ecological experiment images by collecting and annotating the life ecological experiment image products of China’s space station (LEEoCSSCc). Furthermore, we propose an innovative image captioning approach by exploiting cross-modal prediction and data augmentation (IACPDA), which is designed to harness the raw image to guide the semantic constraints of the generated captions from different perspectives. IACPDA facilitates the transformation of both the original image and the resulting caption into a unified semantic space. It then leverages the predictions derived from the original image to offer constructive guidance for the generated captions. Moreover, we use a semantic-guided image augmentation method based on a diffusion model to generate augmented images that are distinct from the original training images while maintaining the essential semantics of the original images. Consequently, these augmented images can be effectively utilized as soft labels, enabling the learning of a more effective generator. The experiments indicate that our method achieves superior performance compared to other state-of-the-art baselines on the LEEoCSSCc. This paper is the first application of image captioning to the analysis of space science experiment images from China’s space station and further offers vital technical support for the understanding of space science experiments.