<p>Tuberculosis (TB) remains a critical global health issue, requiring early and precise diagnosis to reduce transmission and mortality. Chest X-rays (CXRs) are a primary diagnostic tool, but manual interpretation is prone to inconsistencies, necessitating automated solutions. Traditional deep learning approaches, particularly Convolutional Neural Network (CNN)-based models like ResNet and DenseNet, extract local features through hierarchical convolutional layers. However, these models struggle with long-range dependencies, loss of spatial information due to pooling layers, and poor generalization on limited medical datasets. To address these challenges, we propose a Vision Transformer (ViT)-based model enhanced with a Convolutional Stem, Positional Encoding Generator (PEG), and Contrast Limited Adaptive Histogram Equalization (CLAHE). Unlike conventional CNNs, ViTs employ self-attention mechanisms to capture global dependencies, while the convolutional stem aids in low-level feature extraction. PEG further compensates for the lack of spatial inductive bias in ViTs, ensuring better retention of structural information. Additionally, CLAHE improves the contrast of CXRs, making TB-related abnormalities more distinguishable and enhancing feature extraction. Our proposed model achieves superior performance, with 99.04% validation accuracy, 99.28% test accuracy, and a recall of 96.97% for TB cases, ensuring fewer false negatives—critical for clinical applications. Unlike existing CNN- or hybrid-based models, our work uniquely combines CLAHE-based preprocessing, convolutional stems, and positional encoding within a Vision Transformer framework. This integration specifically addresses contrast issues in CXRs and spatial awareness limitations in ViTs, making the approach novel and particularly effective for TB detection. Our study therefore contributes a new, clinically relevant direction in AI-assisted medical imaging.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tuberculosis Detection in Chest X-rays Using Vision Transformers with Convolutional Stems and CLAHE-Based Contrast Enhancement

  • Shanmugavalli Venkatachalam,
  • Chandrasekar Venkatachalam,
  • Priyanka Shah

摘要

Tuberculosis (TB) remains a critical global health issue, requiring early and precise diagnosis to reduce transmission and mortality. Chest X-rays (CXRs) are a primary diagnostic tool, but manual interpretation is prone to inconsistencies, necessitating automated solutions. Traditional deep learning approaches, particularly Convolutional Neural Network (CNN)-based models like ResNet and DenseNet, extract local features through hierarchical convolutional layers. However, these models struggle with long-range dependencies, loss of spatial information due to pooling layers, and poor generalization on limited medical datasets. To address these challenges, we propose a Vision Transformer (ViT)-based model enhanced with a Convolutional Stem, Positional Encoding Generator (PEG), and Contrast Limited Adaptive Histogram Equalization (CLAHE). Unlike conventional CNNs, ViTs employ self-attention mechanisms to capture global dependencies, while the convolutional stem aids in low-level feature extraction. PEG further compensates for the lack of spatial inductive bias in ViTs, ensuring better retention of structural information. Additionally, CLAHE improves the contrast of CXRs, making TB-related abnormalities more distinguishable and enhancing feature extraction. Our proposed model achieves superior performance, with 99.04% validation accuracy, 99.28% test accuracy, and a recall of 96.97% for TB cases, ensuring fewer false negatives—critical for clinical applications. Unlike existing CNN- or hybrid-based models, our work uniquely combines CLAHE-based preprocessing, convolutional stems, and positional encoding within a Vision Transformer framework. This integration specifically addresses contrast issues in CXRs and spatial awareness limitations in ViTs, making the approach novel and particularly effective for TB detection. Our study therefore contributes a new, clinically relevant direction in AI-assisted medical imaging.