The rapid growth of Internet-based marketing on various web platforms has necessitated efficient and cost-effective advertising. There is a lack of comprehensive advertisement generation datasets for training video-generation models. Current datasets consist of simple text captions and videos that lack variety. Additionally, they fail to capture situational backdrops. Hence, to address these challenges, a multimodal dataset, ADGen, was introduced, which comprises 720 variable-length videos totaling 24 h, along with corresponding images and text captions across 15 categories, including Entertainment, Education, and Sports. In addition, a novel multimodal video captioning pipeline is proposed. The Gemini language model was employed to integrate diverse information from video metadata, image descriptions, and audio transcripts for captioning. To assess their strengths and weaknesses, six diverse evaluation metrics were considered, which were validated against human descriptions. Notably, achieving a high BertScore of 0.88 underscores its robust semantic alignment with human references.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ADGen: A Multimodal Advertisement Generation Dataset and Video Captioning Framework

  • Y. P. Pragathi,
  • Swaroop S. Bharadwaj,
  • Preetham Kolli,
  • S. Sanjith,
  • Kavitha Sooda,
  • S. Sunayana

摘要

The rapid growth of Internet-based marketing on various web platforms has necessitated efficient and cost-effective advertising. There is a lack of comprehensive advertisement generation datasets for training video-generation models. Current datasets consist of simple text captions and videos that lack variety. Additionally, they fail to capture situational backdrops. Hence, to address these challenges, a multimodal dataset, ADGen, was introduced, which comprises 720 variable-length videos totaling 24 h, along with corresponding images and text captions across 15 categories, including Entertainment, Education, and Sports. In addition, a novel multimodal video captioning pipeline is proposed. The Gemini language model was employed to integrate diverse information from video metadata, image descriptions, and audio transcripts for captioning. To assess their strengths and weaknesses, six diverse evaluation metrics were considered, which were validated against human descriptions. Notably, achieving a high BertScore of 0.88 underscores its robust semantic alignment with human references.