<p>An accurate and comprehensive semantic understanding, cross-modal reasoning and reliable anomaly detection are necessary for a smart surveillance system to enable real-time decision making. Brother supervised existing deep learning methods are, however, not well suited for semantic alignment of visual and textual information, for unstable video generation and limited robustness in dynamic surveillance environments. These challenges are tackled by this paper, which introduces Quantum-Enhanced Semantic Video Intelligence (QESVI), an integrated system comprising two major components: Quantum-Improved Multimodal Semantic Representation using Contrastive Language–Image Pretraining (CLIP) and Quantum-Inspired Semantic Feature Optimization based and enhanced by Hamiltonian Quantum Generative Adversarial Networks (HQGANs) for video generation. The framework effectively acts as a single structure to support semantic video retrieval, multimodal fusion, anomaly detection and interpretation of the surveillance events. Their proposed method was tested on the UCF-Crime and VIRAT surveillance datasets and compared with recent surveillance-oriented baseline methods, such as SVIP, EAMVF, and UMTAS, using common benchmark metrics, including Semantic Retrieval Precision, Fréchet Video Distance (FVD), Learned Perceptual Image Patch Similarity (LPIPS), Area Under the Curve (AUC), F1-score, False Alarm Rate (FAR), and Frames Per Second (FPS). Experimental results show high performance of QESVI in semantic video understanding with 87% precision in semantic retrieval, 70 fps throughput in the inference phase, an FVD of 328, a motion drift of 0.16, and an anomaly detection confidence of 89%, which reflects the competitive performance of semantic video understanding algorithms for real-time video surveillance. The obtained results indicate that the fusion of CLIP semantic grounding with QIO for tamper-resistant surveillance systems can be effective in enhancing surveillance intelligence in view of the conditions used for the evaluation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Quantum-enhanced Semantic Video Intelligence for Real-time Surveillance Using CLIP and Hamiltonian Quantum GANs

  • S. Saranya,
  • P. Vinayagam

摘要

An accurate and comprehensive semantic understanding, cross-modal reasoning and reliable anomaly detection are necessary for a smart surveillance system to enable real-time decision making. Brother supervised existing deep learning methods are, however, not well suited for semantic alignment of visual and textual information, for unstable video generation and limited robustness in dynamic surveillance environments. These challenges are tackled by this paper, which introduces Quantum-Enhanced Semantic Video Intelligence (QESVI), an integrated system comprising two major components: Quantum-Improved Multimodal Semantic Representation using Contrastive Language–Image Pretraining (CLIP) and Quantum-Inspired Semantic Feature Optimization based and enhanced by Hamiltonian Quantum Generative Adversarial Networks (HQGANs) for video generation. The framework effectively acts as a single structure to support semantic video retrieval, multimodal fusion, anomaly detection and interpretation of the surveillance events. Their proposed method was tested on the UCF-Crime and VIRAT surveillance datasets and compared with recent surveillance-oriented baseline methods, such as SVIP, EAMVF, and UMTAS, using common benchmark metrics, including Semantic Retrieval Precision, Fréchet Video Distance (FVD), Learned Perceptual Image Patch Similarity (LPIPS), Area Under the Curve (AUC), F1-score, False Alarm Rate (FAR), and Frames Per Second (FPS). Experimental results show high performance of QESVI in semantic video understanding with 87% precision in semantic retrieval, 70 fps throughput in the inference phase, an FVD of 328, a motion drift of 0.16, and an anomaly detection confidence of 89%, which reflects the competitive performance of semantic video understanding algorithms for real-time video surveillance. The obtained results indicate that the fusion of CLIP semantic grounding with QIO for tamper-resistant surveillance systems can be effective in enhancing surveillance intelligence in view of the conditions used for the evaluation.