Evaluating Mammogram Image Classification: Impact of Model Architectures, Pretraining, and Finetuning
摘要
This study conducts a thorough evaluation of deep learning architectures, pretraining methods, and finetuning approaches for mammogram classification for tissue density. No architecture was distinctly superior. However, models pretrained on ImageNet consistently surpassed those trained on custom mammogram datasets. Finetuning strategies played a crucial role in model performance. In particular, finetuning the entire model yielded better results. Investigation of confusion matrices revealed that most misclassifications occurred within a one-grade difference, but severe misclassifications were observed in certain configurations. While some architectures offered comparable performance, trade-offs between model performance and computational efficiency were observed, with convolutional neural networks showing faster inference times on CPUs compared to vision transformers.