Audio Driven Video Filtering Using Machine Learning
摘要
Videos typically comprise large volumes of data. The embedding of audio into a video adds to the complexity of analyzing the resulting multimodal data. The user is often interested in only selected components of these recordings. This paper describes the results of our initial research on filtering techniques for such data based on user preferences. Only selected scenes that contain the user specified dialogues and a given set of objects (indicated by the user) are filtered out. The reduction in data volume can significantly reduce the processing time for applications that use filtered data.