Evaluating Deep CNNs and Vision Transformers for Plant Leaf Disease Classification
摘要
The foundation of each nation’s economy has always been agriculture and related sectors. Smart Agriculture is the most recent hot research topic because of its usefulness in different applications, such as early plant disease identification and diagnosis and treatment. Convolutional neural networks (CNNs) have become the de-facto standard in plant leaf disease identification tasks because of their ability to learn complex features. However, not long ago, they began to set new trends in vision tasks alongside the success of the transformer in natural language processing (NLP). This study explores and compares CNNs and Vision Transformer (ViT) models used in the identification of three specific types of plant leaf diseases: Rice Leaf, Tea Leaf, and Maize Leaf images. Their performance is examined using three standard plant leaf image datasets. The study reveals that for all three datasets, Vision Transformer (ViT) outperforms CNNs in terms of the classification of plant leaf diseases. Specifically, the ViT-30 model achieved an average accuracy of 98.41% and 96.95% on Rice Leaf Dataset and Maize leaf dataset respectively, while ViT-20 model achieved an average accuracy of 67.75% on Tea leaf dataset. The main parameters of ViT, such as the optimizer, learning rate, patch size, number of heads, and number of transformer layers, are also fine-tuned, and the optimal ViT configuration for plant leaf disease identification is determined.