Leveraging Large Image-Caption Datasets for Multimodal Taxon Classification
摘要
Taxonomic classification is a fundamental aspect of biology and conservation that poses significant challenges due to the necessity for efficient cross-referencing across a vast taxonomic database. This study explores the efficacy of the CLIP model in enhancing classification across taxonomic ranks by assembling a comprehensive image-caption dataset and aggregating features related to taxonomic hierarchy. The Wikimedia Animals dataset, which consists of approximately 203,000 species image-caption pairs and an average of 1–3 images per species, was used to create representations relevant to animal taxonomy and fine-tune the model on a range of hyper-parameters. Our evaluation reveals divergent model performance along distinct taxonomic rank classifications, with the model trained on a compressed representation of classes demonstrating the highest generalization capability, particularly in the Phylum and Class ranks. Our results provide novel insights into the application of multimodal models in taxonomic classification and highlight potential directions for future research in this field. The development of the large image-caption dataset serves as a benchmark to design models that enhance generalizability for taxonomic classification tasks.