错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Large Image-Caption Datasets for Multimodal Taxon Classification

  • Raynor Kirkson E. Chavez,
  • Kyle Gabriel M. Reynoso,
  • Carlo R. Raquel,
  • Prospero C. Naval

摘要

Taxonomic classification is a fundamental aspect of biology and conservation that poses significant challenges due to the necessity for efficient cross-referencing across a vast taxonomic database. This study explores the efficacy of the CLIP model in enhancing classification across taxonomic ranks by assembling a comprehensive image-caption dataset and aggregating features related to taxonomic hierarchy. The Wikimedia Animals dataset, which consists of approximately 203,000 species image-caption pairs and an average of 1–3 images per species, was used to create representations relevant to animal taxonomy and fine-tune the model on a range of hyper-parameters. Our evaluation reveals divergent model performance along distinct taxonomic rank classifications, with the model trained on a compressed representation of classes demonstrating the highest generalization capability, particularly in the Phylum and Class ranks. Our results provide novel insights into the application of multimodal models in taxonomic classification and highlight potential directions for future research in this field. The development of the large image-caption dataset serves as a benchmark to design models that enhance generalizability for taxonomic classification tasks.