FP-Elegans M1: Feature Pyramid Reservoir Connectome Transformers and Multi-backbone Feature Extractors for MEDMNIST2D-V2
摘要
In medical imaging, several deep learning models, such as vision transformers (ViT), have shown improved capability in recognizing image patterns efficiently by enhancing model efficiency in parameter optimization and sample effectiveness. Our research introduces a fusion approach that leverages multi-backbone pre-trained models (ResNet, EfficientNet, VGG) as feature extractors and the Caenorhabditis elegans’ pyramid connectome ViT in the tail. In most cases, the proposed model demonstrates superior performance on MEDMNIST2D V2.0 classification challenges. Moreover, our approach maintains comparable performance while utilizing fewer training parameters than conventional state-of-the-art models. This indicates a significant step toward more efficient deep-learning architectures in medical diagnostics.