PulmoNetX: A Hybrid Vision Transformer Approach for Multi-scale Spatial Feature Reduction in Pneumonia Classification
摘要
An innovative deep learning structure, PulmoNetX, integrates the capabilities of Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) to enhance pneumonia detection in chest X-ray imagery. During preprocessing, images are normalized in size, converted to grayscale, and subjected to contrast amplification to emphasize essential features. PulmoNetX employs a hybrid methodology to capture both the local and global characteristics of images, leading to significant advancements in diagnosing different pneumonia types, such as COVID-19-induced, viral, and bacterial pneumonia. Comparative studies reveal that PulmoNetX surpasses leading Vision Transformer models in terms of precision, recall, F1-score, and overall accuracy, highlighting its advanced processing abilities and its promise as an effective diagnostic tool in X-ray lung disease detection.