Current Advances in Locality-Based and Feature-Based Transformers: A Review
摘要
The rise of deep learning and the transformative impact of transformers, particularly in NLP and CV domains, have been remarkable. While transformers have shown promise in medical imaging tasks, the original architecture required enhancements to capture local information effectively. Recent research has focused on developing locality-based and feature-based transformers to process spatial information better and improve feature representation in input data. The main content of this survey article includes (1) the paper covers the background of transformers, (2) provides a theoretical review of the transformer and vision transformer, (3) offers insights into different spatial- and feature-based vision transformers, and (4) it also addresses common challenges and explores potential research directions, underscoring the current research challenges in further improving their performance.