QViLa: Quantum Infused Vision-Language Model for Enhanced Multimodal Understanding
摘要
Vision-language models have emerged as transformative tools, revolutionizing the integration of visual and textual information, forging pathways for nuanced interpretations across various applications. The evolution of these models underscores the challenge of achieving seamless modality fusion, particularly in aligning raw pixel values of images with high-level semantics of text, a bottleneck that often hinders optimal cross-modal representations. Addressing this, our research introduces a quantum-enhanced multimodal framework. The main objective of our proposed work is the integration of classical vision-language transformers with quantum-augmented layer, aimed at enhancing the fusion of extracted feature embeddings, thereby bridging the modality gap. The quantum computing techniques, offers an innovative approach to information processing, which paves the way for richer and more intricate visual-textual representations. Furthermore, the shared self-attention mechanism accentuates the model’s ability to detect complex modality interactions. The quantum-enhanced framework is empirically evaluated on the VQA v2 dataset. This evaluation not only considers accuracy across diverse question categories but also the model’s computational efficiency, emphasizing the pivotal contributions of quantum computations in achieving heightened accuracy levels. Further exploration into the influence of different quantum feature maps aided in identifying the most optimal model variant. Our findings highlight the quantum layer’s pivotal role in improving the efficacy of classical vision language models.