NAMN: Normalization-based attention multimodel network for diabetic macular edema grading from multicolor image
摘要
Automated diabetic macular edema (DME) grading technology possesses noteworthy clinical implications, as it facilitates the expeditious and early diagnostic capabilities of ophthalmologists. Existing research leverages fundus imagery as an input to derive grading outcomes. However, the task’s inherent intricacy often results in suboptimal accuracy. This research proposes a normalization based attention multimodal network (NAMN) to improve grading efficiency and accuracy without resorting to lesion segmentation or oversimplifying the grading procedure. NAMN is designed to extract features, incorporating an attention mechanism aimed at more accurately capturing DME specific information within the context of MultiColor image (MCI). The network systematically extracts diverse features from the MCI using a Visual Transformer based multimodal network. Furthermore, the NAM is presented with the aim of diminishing the influence of less significant weights, thereby enabling the model to concentrate on the more critical features within the data.The empirical performance evaluation of the proposed algorithm is conducted using our internally compiled dataset. The classifier exhibits a commendable ability to predict the DME status of MCIs, achieving an accuracy of 0.967, a sensitivity of 0.963, a specificity of 0.939, and an area under the curve (AUC) value of 0.956. The experimental outcomes confirm that the proposed methodology consistently surpasses the performance of the current leading approaches.