FashionGPT: A Large Vision-Language Model for Enhancing Fashion Understanding
摘要
Fashion understanding is a challenging multi-modal task of interpreting multi aspects of fashion images. While traditional computer vision or multi-modal algorithms fall short in providing a comprehensive understanding, Large Vision-Language Model (LVLM) offers a new approach. However, directly using LVLMs presents four major limitations, highlighting the need for a fashion-specific LVLM. Existing fashion datasets also reveal limitations in providing a coherent natural input that fits the LVLMs. To address this bottleneck, we introduce the FUND dataset featuring meticulously annotated textual descriptions for fashion images. Specifically, we build a fashion knowledge base and collect fashion images in various categories online. By leveraging image segmentation model and GPT4, we refine the pre-annotations through manual modifications. Through instruct-tuning with FUND, we develop FashionGPT, a GPT-assisted LVLM based on a solid architecture with exceptional performance on fashion understanding. It is capable of generating coherent and multi-aspect descriptions for fashion images and greatly alleviates the four limitations. Extensive experiments quantitatively and qualitatively demonstrate the effectiveness of FashionGPT and the benefits of FUND, and showcase the broad applications in more tasks.