<p>Visual Polysemy Disambiguation (VPD) and Visual Word Sense Disambiguation (VWSD) are challenging tasks for both computer vision and NLP since an image can have diverse contextual interpretations, ranging from visual representations to abstract concepts. In this paper, we propose a novel approach to address the challenges of VPD and VWSD by leveraging ensemble deep models from computer vision to alleviate the problem of VPD and using the strength of LLMs to mitigate the problem of WSD. We first generate visually representative images from textual descriptions through a zero-shot text-to-image generation framework using image scrapping and Google search. We then employ an ensemble of classifiers and a deep network to learn feature representations, classify images into contexts and find the best match. Similarly, we perform the reverse process of generating textual descriptions from images using a Vision Transformer model and calculate the cosine distance with the actual text. Experimental evaluation of benchmark datasets demonstrates the effectiveness of our combined approach in strengthening both text-to-image and image-to-text generation, we improve disambiguation accuracy, providing a robust solution for VWSD with an MRR of <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11042_2024_20235_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="51" /> </InlineMediaObject> <EquationSource Format="TEX">\(95.77\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>95.77</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> and a Hit rate of <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11042_2024_20235_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="51" /> </InlineMediaObject> <EquationSource Format="TEX">\(92.00\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>92.00</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> surpassing state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging ensemble deep models and llm for visual polysemy and word sense disambiguation

  • Insaf Setitra,
  • Praboda Rajapaksha,
  • Aung Kaung Myat,
  • Noel Crespi

摘要

Visual Polysemy Disambiguation (VPD) and Visual Word Sense Disambiguation (VWSD) are challenging tasks for both computer vision and NLP since an image can have diverse contextual interpretations, ranging from visual representations to abstract concepts. In this paper, we propose a novel approach to address the challenges of VPD and VWSD by leveraging ensemble deep models from computer vision to alleviate the problem of VPD and using the strength of LLMs to mitigate the problem of WSD. We first generate visually representative images from textual descriptions through a zero-shot text-to-image generation framework using image scrapping and Google search. We then employ an ensemble of classifiers and a deep network to learn feature representations, classify images into contexts and find the best match. Similarly, we perform the reverse process of generating textual descriptions from images using a Vision Transformer model and calculate the cosine distance with the actual text. Experimental evaluation of benchmark datasets demonstrates the effectiveness of our combined approach in strengthening both text-to-image and image-to-text generation, we improve disambiguation accuracy, providing a robust solution for VWSD with an MRR of \(95.77\%\) 95.77 % and a Hit rate of \(92.00\%\) 92.00 % surpassing state-of-the-art methods.