<p>Motion prediction in video processing plays a crucial role in action recognition, surveillance, and human–computer interaction. This paper presents a novel framework integrating a Graph Convolutional Network-Inflated 3D Convolution Fusion Network (GCN-I3D FusionNet) with Geese Fused Optimization (GFO) to enhance motion prediction and action recognition. GCN-I3D FusionNet leverages graph-based learning to model complex motion interactions while capturing spatiotemporal depth from video sequences through I3D. Including GFO ensures optimal feature fusion by minimizing redundancy and selecting the most relevant attributes, significantly improving accuracy and efficiency. The proposed method achieves state-of-the-art performance, with accuracy levels of 99.35% on UCF101, 98.58% on JHMDB, and 98.52% on UCF Sports, consistently outperforming the existing models, such as Deep BiLSTM, QSVM, and TSN by a margin of 1.5%–3.2%. Extensive evaluations, including correlation analysis, ROC curves, precision–recall metrics, and F1-score comparisons, further validate the robustness and adaptability of the approach across diverse datasets. These results demonstrate that the projected framework enhances action recognition performance and sets a strong foundation for future research in real-time and scalable video processing systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust motion prediction and recognition using GCN-I3D FusionNet model with optimized feature integration

  • Salini Abraham,
  • T. kanimozhi

摘要

Motion prediction in video processing plays a crucial role in action recognition, surveillance, and human–computer interaction. This paper presents a novel framework integrating a Graph Convolutional Network-Inflated 3D Convolution Fusion Network (GCN-I3D FusionNet) with Geese Fused Optimization (GFO) to enhance motion prediction and action recognition. GCN-I3D FusionNet leverages graph-based learning to model complex motion interactions while capturing spatiotemporal depth from video sequences through I3D. Including GFO ensures optimal feature fusion by minimizing redundancy and selecting the most relevant attributes, significantly improving accuracy and efficiency. The proposed method achieves state-of-the-art performance, with accuracy levels of 99.35% on UCF101, 98.58% on JHMDB, and 98.52% on UCF Sports, consistently outperforming the existing models, such as Deep BiLSTM, QSVM, and TSN by a margin of 1.5%–3.2%. Extensive evaluations, including correlation analysis, ROC curves, precision–recall metrics, and F1-score comparisons, further validate the robustness and adaptability of the approach across diverse datasets. These results demonstrate that the projected framework enhances action recognition performance and sets a strong foundation for future research in real-time and scalable video processing systems.