Caught in the Lens: Zero-Shot BioCLIP Struggles with Camera Trap Wildlife Images
摘要
The decline of wildlife populations and rising species extinctions is an increasing threat do global biodiversity. Camera traps are vital for monitoring animals in their natural habitats. However, manually labeling the resulting images is time-consuming and expensive. This study evaluates the performance of BioCLIP, a specialized Vision Language Model (VLM) for biological images, in classifying camera trap images using zero-shot predictions. We utilized a private dataset from Emirates Nature-WWF, comprising of 65,919 images from the Hajar Mountains in the United Arab Emirates, categorized into 15 classes with significant class imbalance and a dominant “Ghost” class representing non-animal triggers. Using accuracy and F1-score metrics, the zero-shot analysis revealed limited effectiveness, with genus-level prediction accuracy at only 18%, and family and order accuracies at 36% and 42%, respectively. Restricting outputs to the 14 animal genera improved genus accuracy to 38%, but this remains insufficient for practical use. These results indicate that BioCLIP struggles with camera trap images when used directly, highlighting limitations of large vision models trained on idealistic datasets when applied to real-world problems. We recommend fine-tuning or few-shot training to optimize the model for ecological monitoring, suggesting that incorporating object detection could enhance performance and support biodiversity conservation efforts.