A Unified Framework for Few-Shot Medical Image Classification via Multi-agent Description Generation and Refined Contrastive Learning
摘要
Few-shot medical image classification using vision-language models like CLIP is hampered by the semantic gap between coarse labels and fine-grained visual details, limited robustness in standard contrastive learning, and poor interpretability. To address these challenges, we propose a unified framework that integrates a novel multi-agent system with refined contrastive learning objectives for effective CLIP fine-tuning. Our multi-agent system comprises an Attributes Generation Agent (AGA) that mines medical literature for evidence-backed diagnostic attributes and a Captions Generation Agent (CGA) that synthesizes these attributes with image content to generate fine-grained, clinically relevant descriptions. We further introduce a Category-level Contrastive Loss for robust intra-class alignment using these potentially partial descriptions and a Centroid Alignment Loss to enhance feature compactness by aligning image embeddings with stable text centroids. Extensive experiments on four diverse medical imaging datasets demonstrate that our complete framework significantly outperforms state-of-the-art methods, particularly in challenging low-shot scenarios (1–8 shots).