On Using Computer Linguistic Models in the Classification of Biomedical Images
摘要
Computer linguistic models have become widespread in the field of natural language processing (NLP) and have recently been actively used to solve various computer vision problems. In this article, computer studies are carried out aimed at identifying the effectiveness of the use of transformer models in the task of classifying X-ray images of the lungs. The studies use pretrained models of transformers with different sizes ViT-B(16/32) and ViT-L(16/32), which were then fine-tuned on a set of X-ray images of the lungs. Computer studies of the use of the VGG-16, Inception V3, ResNet50, EfficientNetV2, and DenseNet121 convolutional neural networks (CNNs) are also conducted. A comparative analysis of the classification results of the studied X-ray images show that the ViT-B/32 transformer model has the best accuracy metrics with an accuracy of 97.56% and AUC of 99%.