错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empowering few-shot learning: a multimodal optimization framework

  • Liriam Enamoto,
  • Geraldo Pereira Rocha Filho,
  • Li Weigang

摘要

The development of transformer-based models has significantly advanced research in natural language processing and computer vision, allowing us to create models with excellent results across various domains. However, in real-world scenarios, the model may lack generalization ability and perform poorly due to data distribution shifts, insufficient training data, or low-quality data. This work proposes the generic multimodal optimization-based few-shot learning framework (GoFSL). The framework leverages few-shot learning to learn from a few data, multimodal learning to learn a rich representation of image and text data, and meta-learning to help the model generalization. We evaluated the framework using ten datasets from various domains and characteristics, including short texts from Twitter, legal domain long text, text with alphabetic (English and Portuguese) and non-alphabetic (Japanese) languages, images from the medical domain, and multimodal benchmark datasets. GoFSL outperformed the state-of-the-art model ALMO by 1.05% with CUB-200-2011 and multimodal ProtoNet by 0.86% with Oxford-102 dataset. GoFSL is a small but efficient model, with low CO2 estimated emissions (0.01 kgCO2eq), and adaptable to different domains, data modalities, and languages.