Performance of a Convolutional Neural Network Versus Human in Classification of Oral Lesion Images
摘要
This study aims to compare the performance of a convolutional neural network (CNN) with experts in classifying heterogeneous images of oral lesions into six categories of elementary lesions. Additionally, it seeks to determine whether the image characteristics that favor automated classification align with those that result in optimal human performance. A comparative test was conducted between the CNN and human experts, using the same set of images that were employed for CNN training. Additionally, two further tests were conducted to identify the image characteristics that optimize human performance. In the first round, the classification evaluated the area of interest; the experts correctly classified 47.5% of the images, while the CNN correctly classified 77.5%. In the second round, the experts evaluated the entire image with the lesion highlighted by an arrow, with a 57.5% accuracy rate. In the third round, with the lesions segmented, the experts achieved a 70.8% accuracy rate. When comparing the experts’ performance with the CNN, using the same image pattern used in the development of the CNN, only one expert achieved higher accuracy than the CNN: 78.3% and a kappa coefficient of 72.3%. The CNN achieved an accuracy of 77.6% with a kappa coefficient of 71.3%. This study highlights the need to establish specific image characteristics that optimize AI performance in advance. These characteristics differ from those considered ideal for human evaluation. Thus, to implement an AI model in clinical practice with reliable performance, the characteristics of the input images must be carefully defined and standardized.