<p>Accurate classification of brain tumors from magnetic resonance imaging (MRI) remains a challenging task due to the inherent heterogeneity of tumor morphology, class imbalance within datasets, and the limitations of individual deep learning models. To address these challenges, we propose ED-ViTTL (Ensembled Deep Vision Transformer and Transfer Learning), a hybrid framework that leverages both local and global feature representations to enhance diagnostic performance. The model integrates five advanced variants of the Vision Transformer (R50-ViT-L/16, ViT-L/16, ViT-L/32, ViT-B/16, and ViT-B/32) alongside a transfer-learned VGG19 convolutional neural network. Feature embeddings extracted from the ViT and CNN branches are fused through fully connected layers, enabling robust classification into four categories: glioma, meningioma, pituitary tumor, and healthy brain. Experiments were conducted on a publicly available dataset comprising 3264 MRI scans, partitioned into training (70%), validation (15%), and testing (15%) sets using stratified sampling. To mitigate class imbalance and improve model generalization, we employed stratified 5-fold cross-validation, class-weighted categorical cross-entropy, and extensive data augmentation. The best-performing ensemble configuration (ViT-B/32 + VGG19) achieved a classification accuracy of 98.67%, with class-specific AUC values exceeding 0.99 and ROC curves demonstrating clear inter-class separability. Performance metrics, including precision, recall, and F1-scores, remained consistently high across folds, with the confusion matrix indicating minimal misclassifications. These findings demonstrate that ED-ViTTL delivers stable and reproducible results, underscoring its potential as a reliable computer-aided diagnostic tool for assessing brain tumors in clinical practice.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ED-ViTTL: Ensemble Vision Transformer and Transfer Learning Approach for Brain Tumor Classification

  • Amit Thakur,
  • Pawan Kumar Patnaik,
  • Manoj Kumar,
  • Chaitali Choudhary

摘要

Accurate classification of brain tumors from magnetic resonance imaging (MRI) remains a challenging task due to the inherent heterogeneity of tumor morphology, class imbalance within datasets, and the limitations of individual deep learning models. To address these challenges, we propose ED-ViTTL (Ensembled Deep Vision Transformer and Transfer Learning), a hybrid framework that leverages both local and global feature representations to enhance diagnostic performance. The model integrates five advanced variants of the Vision Transformer (R50-ViT-L/16, ViT-L/16, ViT-L/32, ViT-B/16, and ViT-B/32) alongside a transfer-learned VGG19 convolutional neural network. Feature embeddings extracted from the ViT and CNN branches are fused through fully connected layers, enabling robust classification into four categories: glioma, meningioma, pituitary tumor, and healthy brain. Experiments were conducted on a publicly available dataset comprising 3264 MRI scans, partitioned into training (70%), validation (15%), and testing (15%) sets using stratified sampling. To mitigate class imbalance and improve model generalization, we employed stratified 5-fold cross-validation, class-weighted categorical cross-entropy, and extensive data augmentation. The best-performing ensemble configuration (ViT-B/32 + VGG19) achieved a classification accuracy of 98.67%, with class-specific AUC values exceeding 0.99 and ROC curves demonstrating clear inter-class separability. Performance metrics, including precision, recall, and F1-scores, remained consistently high across folds, with the confusion matrix indicating minimal misclassifications. These findings demonstrate that ED-ViTTL delivers stable and reproducible results, underscoring its potential as a reliable computer-aided diagnostic tool for assessing brain tumors in clinical practice.