In the rapidly evolving field of medical imaging, the precise diagnosis and classification of diseases from complex datasets remain a significant challenge. Traditional methods often struggle to capture the intricate details and broader contexts necessary for accurate analysis. This research identifies the limitations of current convolutional neural networks (CNNs) and vision transformers (ViTs) in processing medical images, where CNNs may overlook global image relationships and ViTs might miss local specificities. To address these challenges, we propose a hybrid model that synergistically combines the local feature extraction capabilities of EfficientNet with the global contextual awareness of vision transformers. We design this integrated approach to harness the strengths of both architectures, thereby enhancing the model’s efficiency in performing detailed and context-aware image analysis. By fusing these technologies, the model comprehensively understands medical images, leading to more accurate and reliable diagnostic outcomes. Our experimental results demonstrate that the hybrid model significantly outperforms existing CNN and ViT models’ accuracy and efficiency. Our model achieved a training accuracy of 99.09% on standard medical imaging datasets and a validation accuracy of 92.21%, with substantial improvements in handling complex image features compared to standalone models. The test results further validate the model’s effectiveness, showcasing its potential to revolutionize medical diagnostics. This paper contributes to the advancement of AI in medical imaging, offering a robust solution that enhances diagnostic precision while maintaining computational feasibility.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating EfficientNet and Vision Transformers for Enhanced Medical Image Classification: A Hybrid Approach

  • Amit Taneja,
  • Shubham Sharma

摘要

In the rapidly evolving field of medical imaging, the precise diagnosis and classification of diseases from complex datasets remain a significant challenge. Traditional methods often struggle to capture the intricate details and broader contexts necessary for accurate analysis. This research identifies the limitations of current convolutional neural networks (CNNs) and vision transformers (ViTs) in processing medical images, where CNNs may overlook global image relationships and ViTs might miss local specificities. To address these challenges, we propose a hybrid model that synergistically combines the local feature extraction capabilities of EfficientNet with the global contextual awareness of vision transformers. We design this integrated approach to harness the strengths of both architectures, thereby enhancing the model’s efficiency in performing detailed and context-aware image analysis. By fusing these technologies, the model comprehensively understands medical images, leading to more accurate and reliable diagnostic outcomes. Our experimental results demonstrate that the hybrid model significantly outperforms existing CNN and ViT models’ accuracy and efficiency. Our model achieved a training accuracy of 99.09% on standard medical imaging datasets and a validation accuracy of 92.21%, with substantial improvements in handling complex image features compared to standalone models. The test results further validate the model’s effectiveness, showcasing its potential to revolutionize medical diagnostics. This paper contributes to the advancement of AI in medical imaging, offering a robust solution that enhances diagnostic precision while maintaining computational feasibility.