GPT Vision Meets Taxonomy: A Comprehensive Evaluation for Biological Image Classification
摘要
This study assesses the proficiency of GPT Vision, a multimodal language model, in the classification of biological images across taxonomic ranks. Utilizing a meticulously curated dataset from A–Z Animals, Wikipedia, and eLife Sciences, the model's performance was analyzed at various levels from Phylum to Species. Our quantitative analysis, supported by Python libraries like Pandas, Matplotlib, Seaborn, SciPy, and Statsmodels, revealed a trend of high accuracy at broader taxonomic categories, with a notable decrease at more specific levels. The overall accuracy peaked at 78.95% for Phylum but dropped to 11.58% for Species. Misclassification and null value analyses were conducted, indicating both the strengths of GPT Vision in identifying certain taxa and its limitations, especially at narrower taxonomic levels. The findings underscore the potential of GPT Vision in biological classification while highlighting the necessity for further model refinement, particularly for lower taxonomic ranks.