From Vision to Vocabulary: A Multimodal Approach to Detect and Track Black Cattle Behaviors
摘要
This paper investigates the potential of recent image-text foundation models for classifying black cattle mounting behavior without fine-tuning. Our approach begins with the detection and tracking of each individual black cattle using deep learning-based, fine-tuned YOLOv9 and Deep OC SORT tracking. Once completed, we employ zero-shot approaches, explicitly utilizing the multi-modal Large Language and Vision Alignment (LLAVA) and Large Language Model Meta AI (LLaMA) models. These models integrate visual and linguistic information seamlessly, enabling us to leverage pre-trained knowledge to analyze and understand black cattle behavior directly from images and accompanying text descriptions. By utilizing zero-shot learning, we can bypass the resource-intensive process of model fine-tuning, making it a highly efficient approach for behavior classification. Our approach highlights the robustness and flexibility of multimodal foundation models like LLAVA and LLaMA in handling complex tasks in the agricultural domain, demonstrating their potential for broader applications without requiring extensive retraining on specific datasets. Through our experiments, we showcase the accuracy and efficiency of this zero-shot multimodal approach, providing valuable insights into black cattle mounting behavior that can enhance livestock management and monitoring practices. We introduced a novel cattle dataset tailored for this purpose. We achieved high detection accuracy with a mAP of 0.9856% using our fine-tuned YOLOv9 model and an average tracking accuracy of 94.79% across five videos. The overall accuracy of our integrated system demonstrates its efficacy in accurately classifying and tracking cattle behaviors, underscoring the potential of zero-shot learning models in precision agriculture.