Animal classification and counting are critical in agriculture, wildlife conservation, and livestock monitoring. This study investigates the use of Vision Transformers (ViT) for automated animal identification and counting from images. We employed a Kaggle dataset containing 5,400 images across 90 animal species. Images were resized to 224 × 224 and normalized to match ViT input requirements. The dataset was split into 70% training, 15% validation, and 15% testing. We fine-tuned the pre-trained ViT model (google/vit-base-patch16–224-in21k), modifying the final layer for 90-class classification. Training was performed over three epochs using the Adam optimizer with a learning rate of 2 × 10−5 and batch size of 16. Model performance was evaluated using accuracy, precision, recall, and F1-score. Animal counts were estimated by analyzing predicted label frequencies. The final model achieved a classification accuracy of 96.36% and a counting accuracy of 98.94%, demonstrating strong performance across tasks. These results highlight the effectiveness of ViT in fine-grained image recognition and its potential for intelligent animal monitoring applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid Spatio-Temporal Transformer with Diffusion-Guided Attention for Accurate Animal Counting and Classification in Dense Environments

  • Kanda Tshinu Patrick,
  • Chunling Tu,
  • Owolawi Pius Adewale,
  • Antonie Smith

摘要

Animal classification and counting are critical in agriculture, wildlife conservation, and livestock monitoring. This study investigates the use of Vision Transformers (ViT) for automated animal identification and counting from images. We employed a Kaggle dataset containing 5,400 images across 90 animal species. Images were resized to 224 × 224 and normalized to match ViT input requirements. The dataset was split into 70% training, 15% validation, and 15% testing. We fine-tuned the pre-trained ViT model (google/vit-base-patch16–224-in21k), modifying the final layer for 90-class classification. Training was performed over three epochs using the Adam optimizer with a learning rate of 2 × 10−5 and batch size of 16. Model performance was evaluated using accuracy, precision, recall, and F1-score. Animal counts were estimated by analyzing predicted label frequencies. The final model achieved a classification accuracy of 96.36% and a counting accuracy of 98.94%, demonstrating strong performance across tasks. These results highlight the effectiveness of ViT in fine-grained image recognition and its potential for intelligent animal monitoring applications.