Self-supervised approach for diabetic retinopathy severity detection using vision transformer
摘要
Diabetic retinopathy (DR) is a diabetic condition that affects vision, despite the great success of supervised learning and Conventional Neural Networks (CNNs), it’s still challenging to detect the severity of DR at an early stage. The label-intensive nature of supervised learning and the limited scalability of CNNs inhibit exploiting tons of unlabeled medical images that can be useful for capturing rich domain-specific features. The local feature representations from CNNs and their inability to scale well with increased unlabeled data lead to indiscriminative representations not effective for the downstream task. Hence in this work, vision transformer-based paradigm for diabetic retinopathy (DR) severity detection is presented by undertaking a scalable learning approach for model building. Self-supervised learning (SSL) framework with transformer architecture is used to classify fundus images to one of the DR categories by pre-training the model on tons of unlabeled fundus images followed by supervised training on different fractions of labeled data. The performance of the proposed transformer based self-supervised DR detection models (