Enhancing interpretability in film shot analysis through continuous shot integration and saliency maps
摘要
In film studies, there is potential value in conducting interpretable analysis of film shot language, as the preceding and following shots of a shot significantly influence its scale, movement, and composition choices. However, previous methods of analyzing film shot attributes have lacked research into the interpretability of shot type analysis results, often focusing solely on individual shots and neglecting the impact of preceding and succeeding shots on the current shot analysis. Therefore, in this study, we propose a new research approach. Specifically, we utilize information from continuous shots to analyze the attributes of the current shot and enhance the model’s interpretability by integrating a model construction based on saliency maps. We design a training framework that takes continuous shots as input and utilizes saliency maps to guide model training. Additionally, we apply masks to frames from the current shot as input and design a consistency module to emphasize the temporal attributes of the shot. Experimental results on our dataset with continuous shots demonstrate that our proposed method outperforms all previous methods.