Visual fixation guided spatio-temporal dual branch network for gaming video quality assessment
摘要
Recently, the application of gaming video has become increasingly widespread, making the evaluation of the gaming video quality an important research topic. Considering that gaming videos contain a large amount of temporal and spatial redundancy, and different frames and regions have significantly different visual attractiveness, this paper innovatively proposes a visual fixation guided dual branch network (VFGDNet) for spatio-temporal sampling and no-reference gaming video quality assessment (NR-GVQA). First, based on the visual fixation principle, a spatio-temporal sampling strategy for gaming video is proposed. In the temporal domain, a visual fixation intensity guided key frame extraction method is presented. In the spatial domain, a visual fixation guided global and local grid dual-spatial sampling method is proposed to effectively extract visually important information from key frames. Furthermore, this paper proposes a dual-branch network containing motion and spatial feature extraction branches for quality prediction. The motion branch uses a pre-trained SlowFast backbone network to extract temporal features from the scaled key frames to express the fluency of the video, while the spatial branch uses a ViT backbone network to extract spatial scene and content features from the global and local grid sampling maps. Finally, the motion and spatial features are fused and the video quality is predicted using a multi-layer perceptron (MLP). Experiments were conducted on four gaming video datasets: GamingVideoSet, KUGVD, LIVE-Meta MCG, and LIVE-YT-Gaming, and the results show that our VFGDNet is superior to existing methods.