错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bilateral Lesion-Guided Transformers with Patient-Level Graph Reasoning for Multi-label Retinal Disease Diagnosis

  • Muhammad Ali Iqbal,
  • Inchul Han,
  • Soo Kyun Kim

摘要

Accurate multi-label recognition of retinal diseases from color fundus photographs is essential for scalable ophthalmic screening. We present Bi–LGT, a bilateral lesion-guided transformer that exploits the natural pairing of left and right eyes. For each eye, a high-resolution CNN with lesion-aware gating produces multi-scale tokens and global descriptors. A bidirectional, token-level cross-attention exchanges information prior to per-label representation aggregation, capturing inter-eye asymmetry and complementarity. Learnable class queries then perform label-specific token aggregation to produce calibrated decision logits, while a graph-regularized patient head propagates information over a learned disease co-occurrence graph. The model is trained end-to-end with focal losses and an eye patient consistency term. On an eight-label fundus benchmark, Bi–LGT achieves uniformly strong discrimination, with per-class AUCs of 0.93–0.98 (macro AUC \(\approx 0.961\) ) and macro F1 \(\approx 0.93\) . In particular, Cataract (F1 = 0.98, AUC = 0.98) and Myopia (F1 = 0.96, AUC = 0.97) are highest, while Normal is most challenging (AUC = 0.93, recall = 0.87), likely reflecting borderline presentations in multi-label screening. Compared with strong CNN baselines, Bi–LGT raises AUC by 7–15 percentage points (98.76 vs. 83.93–91.30), substantially improves inter-rater agreement (Cohen’s \(\kappa =87.92\) ), and yields a higher composite score (Final = 91.41). Qualitative checks show steep ROC behavior and clinically plausible results. Overall, bilateral token exchange, class-query–based label-specific aggregation, and graph-aware inference yield an accurate, interpretable, and deployable solution for comprehensive multi-label screening from fundus images.