ADGen: A Multimodal Advertisement Generation Dataset and Video Captioning Framework
摘要
The rapid growth of Internet-based marketing on various web platforms has necessitated efficient and cost-effective advertising. There is a lack of comprehensive advertisement generation datasets for training video-generation models. Current datasets consist of simple text captions and videos that lack variety. Additionally, they fail to capture situational backdrops. Hence, to address these challenges, a multimodal dataset, ADGen, was introduced, which comprises 720 variable-length videos totaling 24 h, along with corresponding images and text captions across 15 categories, including Entertainment, Education, and Sports. In addition, a novel multimodal video captioning pipeline is proposed. The Gemini language model was employed to integrate diverse information from video metadata, image descriptions, and audio transcripts for captioning. To assess their strengths and weaknesses, six diverse evaluation metrics were considered, which were validated against human descriptions. Notably, achieving a high BertScore of 0.88 underscores its robust semantic alignment with human references.