<p>With the widespread adoption of high-speed networks such as 4G and 5G, along with the explosive growth of social media platforms, video content is now frequently captured and shared online without accompanying metadata such as tags or descriptions. This absence of textual annotations presents a significant challenge for indexing and retrieving relevant video content. Content-Based Video Retrieval systems address this issue by analyzing the visual content of videos rather than relying on external metadata. However, only limited efforts in the literature have jointly explored both the spatial and temporal context of video data for retrieval. To address this gap, we propose a Content-Based Video Retrieval framework that leverages a 3D Convolutional Neural Network, specifically the R(2+1)D architecture enhanced with transfer learning. This model decomposes spatiotemporal convolutions to more effectively capture both spatial and temporal video features. In addition, we introduce a novel classification-similarity-based weighted distance approach, which overcomes the limitations of traditional distance-based and classifier-based retrieval methods. Experimental evaluation on the UCF101 dataset demonstrates that the proposed system achieves a significant improvement in retrieval performance, with over a 20% increase in AUC compared to baseline techniques.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging 3DCNN and Weighted Similarity Metrics for Enhanced Content-Based Video Retrieval

  • Farooq Shaik,
  • Ashu Abdul,
  • Jatindra Kumar Dash

摘要

With the widespread adoption of high-speed networks such as 4G and 5G, along with the explosive growth of social media platforms, video content is now frequently captured and shared online without accompanying metadata such as tags or descriptions. This absence of textual annotations presents a significant challenge for indexing and retrieving relevant video content. Content-Based Video Retrieval systems address this issue by analyzing the visual content of videos rather than relying on external metadata. However, only limited efforts in the literature have jointly explored both the spatial and temporal context of video data for retrieval. To address this gap, we propose a Content-Based Video Retrieval framework that leverages a 3D Convolutional Neural Network, specifically the R(2+1)D architecture enhanced with transfer learning. This model decomposes spatiotemporal convolutions to more effectively capture both spatial and temporal video features. In addition, we introduce a novel classification-similarity-based weighted distance approach, which overcomes the limitations of traditional distance-based and classifier-based retrieval methods. Experimental evaluation on the UCF101 dataset demonstrates that the proposed system achieves a significant improvement in retrieval performance, with over a 20% increase in AUC compared to baseline techniques.