Foundation model-based haemoglobin concentration estimation and anaemia risk assessment from retinal fundus images
摘要
Widespread screening of anaemia is currently limited by the need for invasive blood tests. This study’s primary objective was to evaluate the feasibility and interpretability of Vision Transformer (ViT)-based foundation models, including those pretrained on natural or medical domains, for the noninvasive estimation of haemoglobin (Hb) levels using retinal fundus images.
MethodsThis retrospective cross-sectional study included 34,511 individuals from the Health Promotion Center of Seoul National University Hospital (2005–2016). We adapted five foundation models (DINOv2, OpenCLIP, MAE, RETFound, and VisionFM) for Hb concentration regression using low-rank adaptation (LoRA) and partial fine-tuning. For explainability, we applied Grad-CAM saliency mapping and perturbation-based analysis using retinal vessel segmentation.
ResultsAmong all tested models, the natural-domain model DINOv2 combined with LoRA achieved the best performance (MAE = 0.703, R2 = 0.682, and AUROC = 0.917 for anaemia detection). Operating characteristics were examined at fixed specificity levels across different haemoglobin thresholds. Grad-CAM analysis revealed that the model’s predictions predominantly relied on the macula and peripapillary regions. Perturbation analysis confirmed that increased brightness and colour saturation of retinal arteries correlated with higher estimated Hb levels.
ConclusionsFine-tuned foundation models, particularly DINOv2 with LoRA, demonstrated competitive performance in noninvasive haemoglobin estimation from retinal fundus images. The model’s decision-making process was found to be qualitatively consistent with clinical and physiological expectations, with predictions focusing on known anatomical and vascular indicators. These preliminary findings suggest the potential use of adapted foundation models as adjunctive tools for anaemia risk assessment, though external validation remains necessary.