The video’s brightness, color, shot composition, and camera information are key visual elements that reflect the creator’s photographic style. To address the challenge of systematically extracting complex visual features, we have researched how to automatically extract five types of visual styles: brightness, color, camera stability, editing rhythm, and shot composition, and proposed a multi-level Visual Narrative Network (VNNet) system. To analyze the style of different categories of videos and compare the differences in visual expression techniques among creators, we have constructed a Cinematic Style Dataset. This dataset covers a variety of video categories, allowing us to delve into the visual differences between different works and explore the stylistic differences of individual creators. Experimental results show that VNNet has validated the effectiveness of the tags on the Cinematic Style Dataset through manual analysis, with an overall consistency rate of 75.29%, demonstrating its effectiveness and practicality in video visual style classification. These findings not only provide a new perspective for the automatic analysis and categorization of video content but also offer technical support for future video understanding and content creation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VNNet: A Deep Learning-Based System for Video Visual Style Classification

  • Yaxin Bai,
  • Qinglan Wei,
  • Li Yang

摘要

The video’s brightness, color, shot composition, and camera information are key visual elements that reflect the creator’s photographic style. To address the challenge of systematically extracting complex visual features, we have researched how to automatically extract five types of visual styles: brightness, color, camera stability, editing rhythm, and shot composition, and proposed a multi-level Visual Narrative Network (VNNet) system. To analyze the style of different categories of videos and compare the differences in visual expression techniques among creators, we have constructed a Cinematic Style Dataset. This dataset covers a variety of video categories, allowing us to delve into the visual differences between different works and explore the stylistic differences of individual creators. Experimental results show that VNNet has validated the effectiveness of the tags on the Cinematic Style Dataset through manual analysis, with an overall consistency rate of 75.29%, demonstrating its effectiveness and practicality in video visual style classification. These findings not only provide a new perspective for the automatic analysis and categorization of video content but also offer technical support for future video understanding and content creation.