Beyond Accuracy: A MultiDimensional Framework for Evaluating Medical Image Classification Through Win vs. Lose Model Comparisons
摘要
High-performing deep learning models such as ResNet, originally optimized for large-scale natural image datasets, often fail to generalize when applied directly to medical imaging tasks. This study investigates the limitations of “off-the-shelf” models in the context of skin lesion classification using the DermaMNIST dataset. Through a systematic evaluation of 35 architectural configurations across varying image resolutions and depths, the analysis reveals that mid-depth architectures (3–4 layers) and intermediate resolutions (