<p>Content-based mammographic image retrieval requires exact BIRADS categorical matching across five classes, posing far greater complexity than conventional binary classification. Existing studies remain limited by small sample sizes, improper patient-level separation, and inadequate statistical validation, restricting clinical translation. We developed a comprehensive evaluation framework systematically comparing CNN architectures (DenseNet121, ResNet50, VGG16) under advanced training strategies: fine-tuning, metric learning, and super-ensemble optimization. Rigorous patient-stratified splits (1003 patients, two images each), 602 test queries, and bootstrap confidence intervals (1000 resamples) ensured reliable assessment. Advanced fine-tuning and test-time augmentation (TTA) yielded a precision@10 of 34.71% for DenseNet121 _AdvancedFT_TTA, a 25.74% improvement over the baseline ResNet50 (27.6%). Selective super-ensemble and metric learning approaches were further benchmarked under patient-exclusive splits, confirming robust performance across architectures. Statistical analysis (bootstrap CIs, <i>n</i> = 1000; <i>t</i>-tests <i>p</i> &lt; 0.001; Cohen’s <i>d</i> &gt; 0.8) validated significant gains and reproducibility. These results establish DenseNet121_AdvancedFT_TTA as the new state-of-the-art for five-class BIRADS retrieval while reducing computational cost.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advanced Multi-architecture Deep Learning Framework for BIRADS-Based Mammographic Image Retrieval: Comprehensive Performance Analysis with Super-Ensemble Optimization

  • MD Shaikh Rahman,
  • Feiroz Humayara,
  • Syed Maudud E. Rabbi,
  • Muhammad Mahbubur Rashid

摘要

Content-based mammographic image retrieval requires exact BIRADS categorical matching across five classes, posing far greater complexity than conventional binary classification. Existing studies remain limited by small sample sizes, improper patient-level separation, and inadequate statistical validation, restricting clinical translation. We developed a comprehensive evaluation framework systematically comparing CNN architectures (DenseNet121, ResNet50, VGG16) under advanced training strategies: fine-tuning, metric learning, and super-ensemble optimization. Rigorous patient-stratified splits (1003 patients, two images each), 602 test queries, and bootstrap confidence intervals (1000 resamples) ensured reliable assessment. Advanced fine-tuning and test-time augmentation (TTA) yielded a precision@10 of 34.71% for DenseNet121 _AdvancedFT_TTA, a 25.74% improvement over the baseline ResNet50 (27.6%). Selective super-ensemble and metric learning approaches were further benchmarked under patient-exclusive splits, confirming robust performance across architectures. Statistical analysis (bootstrap CIs, n = 1000; t-tests p < 0.001; Cohen’s d > 0.8) validated significant gains and reproducibility. These results establish DenseNet121_AdvancedFT_TTA as the new state-of-the-art for five-class BIRADS retrieval while reducing computational cost.