Cross-Modal Text-to-Video Retrieval Using Deep Learning
摘要
In the current digital environment, speech, visual signals, and sound play pivotal roles, particularly in online videos with informative audio tracks. The vastness of multimedia content on the Internet presents a unique challenge, one addressed through the application of cutting-edge technologies such as Transformer-based models and cosine similarity. Research in this domain focuses on retrieving videos using natural language queries. Past efforts aimed to bridge this gap, with deep learning offering a promising solution. Platforms like TikTok, YouTube, and Netflix connect global audiences, yet finding semantically similar videos via text queries remains a vital task. Amid this digital revolution, this research introduces an innovative cross-modal attention mechanism meticulously engineered for text-to-video retrieval, transcending conventional approaches. By enabling precise video searches, enhancing search accuracy, and optimizing user experiences within the vast realm of online multimedia resources, this work stands at the forefront of reshaping how users engage with digital video content, adapting to the ever-changing landscape of the digital age.