Optimization of music teaching model based on video image semantic analysis and retrieval
摘要
Video semantic analysis is essential for interpreting audiovisual content, as conventional methods often face challenges such as slow processing, limited accuracy, and poor emotional understanding, hindering effective comprehension and automated content interpretation. The proposed framework utilizes a Spatiotemporal Mode Graph-based Temporal Saliency Mapping (SMGTSM) algorithm to enhance frame-wise saliency map generation and support real-time video analysis. SMGTSM achieves a processing time of 523.7 milliseconds, outperforming optical flow and Random Sample Consensus (RANSAC), while maintaining high interpretative accuracy. The system reached 97% accuracy in music feature recognition and 94.67% accuracy, 95.13% recall, and 93.95% F1-score in emotion classification. The result demonstrated that SMGTSM effectively captures subtle spatiotemporal and emotional patterns in video. By leveraging these capabilities, the approach enables real-time interaction, adaptive music interaction, and fosters increased student engagement. The suggested model bridges theoretical instruction with experiential learning by integrating content with video semantics, highlighting the transformative impact of semantic video analysis in music education.