Emocoop: dynamic vision–language coupling for multi-label emotion classification in Dongba paintings
摘要
Dongba paintings, a unique Naxi cultural heritage, are characterized by complex symbolism and diverse visual aesthetics. However, contemporary vision–language models exhibit insufficient domain generalization capabilities in few-shot emotion classification on these artworks, substantially impeding accurate multi-label emotion recognition. To address this critical challenge, we propose EmoCoOp, a dynamic vision–language coupling framework. It incorporates latent priors from visual semantics to guide prompt learning and employs a meta-network for adaptive visual prompt generation. A vision-language coupling mechanism facilitates deep multimodal integration, while a dual-path chromatic affective inference module models coloration and emotional expressions. Experimental results demonstrate EmoCoOp’s superior performance, achieving 75.73% mAP and 82.21% Recall@2, outperforming the second-ranked model by significant margins. This framework advances multi-label emotion classification for ethnic artworks and cross-modal understanding. The code is available at https://github.com/yang-easy/EmoCoOp.