Self-supervised Multimodal Stance Detection with BERT and ResNet
摘要
Presenting a modern strategy for stance detection that leverages self-supervised learning with transformer-based designs. Our approach to coordinating content and picture information using BERT to encode content and ResNet for pictures includes extraction. The demonstration is to begin with pre-trained with a contrastive loss function to adjust multimodal representations, at that point fine-tuned for stance classification. Typically, it is a step up from traditional strategies that as it were consider content because it incorporates an imperative visual setting. The tests on a benchmark dataset show that the strategy progresses execution, demonstrating the viability of self-supervised multimodal learning. To dig into a point-by-point execution analysis, indicating common mistakes, impediments, and future research opportunities.