Abstract <p>The problem of stability of the classifying neural network model to rotations of the object in the image plane, which is critical in practical applications, is considered. The paper shows that the classification accuracy of the baseline Vision Transformer model drops significantly, by more than 3%, when input images are rotated by 20°. The task is to modify the basic neural network model of image classification to increase its reliability when rotating input images. To quantitatively assess the reliability of the model, it is proposed to use not only the classification accuracy but also the cosine similarity measures and <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11493_2025_8729_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\({{L}_{1}}\)</EquationSource> <!--PatRec2570029Dmitrienko-m1--> </InlineEquation>-distances between feature vectors of the original and rotated images in the latent space of the classifier. An approach is proposed for integrating a spatial–transformation module into the latent space of the base model and training it through parallel branches with pretraining, which allows for a significant increase in the model’s stability to input image rotations up to 20° with a slight decrease in classification accuracy of only 0.2% compared to the accuracy of the baseline model on unrotated images. Further work may generalize the obtained result to a wider class of transformations and neural network architectures.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Rotation Robust Image Classification with Visual Transformer Based on Spatial Transformation of Feature Maps

  • A. Dmitrienko,
  • A. Gneushev

摘要

Abstract

The problem of stability of the classifying neural network model to rotations of the object in the image plane, which is critical in practical applications, is considered. The paper shows that the classification accuracy of the baseline Vision Transformer model drops significantly, by more than 3%, when input images are rotated by 20°. The task is to modify the basic neural network model of image classification to increase its reliability when rotating input images. To quantitatively assess the reliability of the model, it is proposed to use not only the classification accuracy but also the cosine similarity measures and \({{L}_{1}}\) -distances between feature vectors of the original and rotated images in the latent space of the classifier. An approach is proposed for integrating a spatial–transformation module into the latent space of the base model and training it through parallel branches with pretraining, which allows for a significant increase in the model’s stability to input image rotations up to 20° with a slight decrease in classification accuracy of only 0.2% compared to the accuracy of the baseline model on unrotated images. Further work may generalize the obtained result to a wider class of transformations and neural network architectures.