Multimodal Learning for Road Safety Using Vision Transformer ViT
摘要
This paper proposes a novel approach for multimodal learning that combines visual information from images with structured data from a multi-column dataset. The approach leverages the power of vision transformer (ViT) models, which are based on self-attention layers, to extract high-quality visual features from images. The proposed approach was evaluated on a road safety problem, specifically the identification of high-risk areas for road accidents. The experiment was conducted on a dataset of road sections in Morocco and demonstrated the effectiveness of the proposed approach in improving the accuracy of the classification task. The results show that the multimodal approach outperformed unimodal approaches that relied solely on either visual data or structured data. The proposed approach can be applied to various domains beyond road safety, such as healthcare, finance, and social media analysis, where multimodal data is abundant and complex relationships between different data modalities need to be captured. The contribution of this work is the development of a multimodal classification model that combines visual and structured data leveraging the power of vision transformer (ViT) for image encoding, resulting in high-quality visual features. This approach is a departure from classical convolutional models, which lack self-attention layers, making it suitable for a wide range of applications.