A Review of Visual Transformer Research
摘要
The development of Transformer in the field of computer vision has been very rapid in the past two years. Influenced by the development of Transformer in natural language processing and the research ideas of Vision Transformer model using multi-head attention in image classification, more and more researchers are paying attention to the application of Transformer. The application prospects of Transformer models in computer vision are broad, but there are still some inherent problems that need to be overcome. This article introduces the basic principles of Transformer and Vision Transformer, and focuses on the problems that arise when applying Transformer in the field of computer vision, such as large computation and data requirements. From these issues, the author analyzes the improvement strategies proposed in related papers, and finally points out the future development direction of visual Transformers based on relevant models, in order to further promote research in this field and explore more efficient visual Transformer models.