REMEDI: a multimodal deep learning framework for diabetic retinopathy follow-up prediction
摘要
Diabetic Retinopathy (DR) is a major cause of vision loss in diabetic patients, highlighting the need for predictive models of its progression. While deep learning has improved DR detection and grading, temporal and spatial prediction remain underexplored. We propose REMEDI, a multimodal deep learning framework that predicts both clinical evolution and spatial distribution of retinal lesions over time. Using clinical variables (BCVA, DRSS, CST) and the current lesion mask, REMEDI generates the predicted segmentation mask for the next visit, explicitly modeling temporal dynamics of clinical and spatial features. The framework jointly performs three tasks: (i) regression of clinical variables; (ii) prediction of lesion geometry evolution; and (iii) generation of follow-up lesion masks. These modalities are hierarchically integrated: clinical and morphological regression outputs, together with the baseline fundus mask, are used to inform the generation of the disease mask at the follow-up visit, capturing how the treatment is influencing the lesion progression. Results show that the LSTM and MLP models yield the most accurate clinical and morphometric predictions, while standard U-Net and Attention U-Net outperform more complex architectures in temporally consistent segmentation. REMEDI provides temporally consistent and interpretable forecasts, supporting personalized DR monitoring and treatment planning.