A deep learning based automatic report generator for retinal optical coherence tomography images
摘要
Reading and summarizing insights from Optical Coherence Tomography (OCT) images is a routine yet time-consuming task that requires expensive time from experienced ophthalmologists. This paper introduces the Multi-label OCT Report Generation (MORG) model, a deep learning approach to assist in the interpretation of OCT images. MORG employs dual image encoders to extract features from OCT image pairs, fusing them through a multi-scale module with an attention mechanism, followed by a sentence decoder to produce reports. Trained and tested on 57,308 retinal OCT image pairs, MORG achieved high classification accuracy for 16 pathologies with 37 descriptive types. It also excelled in a blind grading test against general large language models and other state-of-the-art image captioning models, scoring 4.55 compared to ophthalmologists’ 4.63 out of a maximum of 5. Furthermore, MORG has the potential to reduce the report drafting time for ophthalmologists by 58.9%, significantly alleviating their workload.