Complementary Deep Features for COVID-19 Screening: Combining Hierarchical Vision Transformers and Convolutional Networks
摘要
Chest X-ray screening tools must be accurate and practical, especially when compute is limited. We fine-tune VGG19 and the Swin Transformer separately and combine their logits in a two-model ensemble. We detail preprocessing, training, and fusion so others can reproduce the setup. On the COVID-19 Radiography dataset, stratified five-fold cross-validation shows the ensemble improves over both backbones in accuracy (99.32% ± 0.19), sensitivity, specificity, and F1. We average weights only within a shared architecture (“model soups”) and average predictions across different architectures.