Gaze-and-Machine Dual-Driven Attention Fusion Network for Medical Image Classification
摘要
In medical image recognition, challenges such as data scarcity, inter-individual variability, and noise often lead to shortcut learning, where models fixate on visual cues irrelevant to the underlying pathology. Human gaze data naturally highlights critical regions and, when combined with the expert knowledge of clinicians, markedly improves diagnostic accuracy. However, the inability to precisely emulate expert knowledge creates a significant gap between human and model cognition, limiting the accuracy of deep learning models that rely solely on sparse gaze data. To address this, we introduce the Gaze-and-Machine Dual-driven Attention Fusion Network (GMD-AFNet), designed to enhance feature extraction in medical image classification. GMD-AFNet integrates raw image data, human eye-tracking information, and model-generated attention within a multi-branch architecture. By leveraging diversity loss, it encourages each branch to capture complementary and unique features, avoiding redundant representations. Contrastive loss enhances feature discriminability and generalization by pulling together features of similar samples and pushing apart those of dissimilar ones, thereby improving classification performance. We evaluate GMD-AFNet on the CXR-Eye dataset, which includes eye-tracking data from radiologists diagnosing chest X-rays, providing a realistic depiction of expert attention patterns. Experimental results demonstrate that GMD-AFNet effectively fuses human and machine attention, achieving substantial improvements in classification accuracy over baseline models.