Agro-LLaVA-Next: A Large Multimodal Model for Plant Diseases Recognization
摘要
This paper presents Agro-LLaVA-Next, a specialized framework aimed at overcoming the limitations of general-purpose large multimodal models (LMMs) in plant disease recognition. We propose a parameter-efficient fine-tuning strategy that freezes the language module while selectively optimizing the visual encoder and multimodal projector. This approach ensures ffective adaptation to agricultural vision tasks without compromising the model’s inherent linguistic capabilities. To address the scarcity of data in agricultural visual question answering, we construct the AgroTest dataset, which includes 73,125 training samples across 22 crop categories contain 117 disease or health types, through strategic data synthesis. Specifically, 2% of the total 74,548 images were randomly and equally stratified sampled to form the test set (1,423 samples), while the remaining 98% were used to generate training data. Single-turn question-answer (QA) pairs for the training set were synthesized using predefined templates, whereas QA pairs for the test set were generated as multiple-choice questions using the Claude 3.5 Sonnet model. Experimental validation on the test set demonstrates a significant 30.2% improvement in disease recognition accuracy compared to baseline LMMs, with only a moderate 10.6% performance degradation in general dialogue tasks. This establishes an effective trade-off for agricultural applications. The proposed framework underscores the potential of targeted LMM adaptation in addressing domain-specific visual recognition challenges.