EfficientPhoCaption: An Enhanced PhoBERT-Based Image Captioning Framework for Breast Cancer Diagnosis
摘要
In an era marked by rapid technological progress, artificial intelligence (AI) has become increasingly prominent in supporting diagnostic and therapeutic practices across various fields of medicine. One notable application is the use of AI to automatically generate descriptions for medical images, a technique that has the potential to streamline diagnostic workflows and significantly reduce costs for patients. In this study, we present EfficientPhoCaption, a novel framework that integrates an advanced deep learning architecture combining EfficientNetB3 for image feature extraction and the PhoBERT base model for tokenizing Vietnamese textual descriptions. To enhance the quality of both images and texts during the preprocessing stage, we also employed several sophisticated techniques, such as Contrast Limited Adaptive Histogram Equalization (CLAHE) and normalization approach. Extensive experiments demonstrate that the proposed method achieves outstanding performance across multiple evaluation metrics. For instance, on the HisBreast dataset, the model achieved a BLEU score of 0.8761—the highest reported for this dataset. Additionally, the model attained impressive average scores of 0.6765 for Rouge-L and 0.5070 for GLEU. Similarly, when evaluated on the INBreast dataset, EfficientPhoCaption delivered remarkable results, achieving BLEU, Rouge-L, and GLEU scores of 0.7223, 0.5509, and 0.3804, respectively. Compared to the previous BCICG (Breast Cancer Image Caption Generator) method, our approach shows substantial improvements. BLEU-4, for example, increased by 149.15%, rising from 0.234 to 0.583. BLEU-1 and BLEU-2 also improved by over 81%, while BLEU-3 saw a 130.24% increase. Rouge-L exhibited a more moderate improvement of 15.33%, and for the first time, GLEU was reported with a strong score of 0.507.