Artificial intelligence versus human expertise: reliability of ChatGPT and the London atlas for dental age estimation using panoramic radiographs
摘要
This study evaluated the performance of ChatGPT, a multimodal large language model (LLM), in estimating dental age from panoramic radiographs (PRs) and compared its accuracy and reproducibility with those of the London Atlas (LA) method.
MethodsPRs of 620 healthy children aged 6 through 13 years were retrospectively analyzed. An experienced dentomaxillofacial radiologist estimated dental age using the LA, and the ChatGPT-4o model analyzed the same anonymized images to generate automated age predictions. Both methods were repeated after two weeks to assess intra-observer reliability. Predictive accuracy and agreement with chronological age (CA) were evaluated using mean absolute error (MAE), root mean squared error (RMSE), intraclass correlation coefficients (ICC), and Bland–Altman analyses. Statistical significance was set at p < .05.
ResultsChatGPT’s predictions differed significantly from chronological age (CA), tending to overestimate age in younger children and underestimate age in older children. Compared with the LA, ChatGPT exhibited higher MAE and RMSE values, indicating lower predictive accuracy and greater variability. Error magnitudes were greatest in the 6-, 12-, and 13-year-old groups and lowest in the 8-year-old group, whereas the LA showed lower and more stable errors across ages. The LA demonstrated fewer discrepancies and excellent reproducibility (ICC = 0.960) as compared with the moderate agreement of ChatGPT (ICC = 0.703). Overall, the LA provided estimates closer to CA, whereas ChatGPT exhibited greater variability.
ConclusionsChatGPT shows promise for complex decision-making tasks such as dental age estimation; however, its current accuracy, reproducibility, and output stability remain inferior to established methods such as the LA. The inconsistent predictions observed across repeated evaluations highlight a critical limitation regarding its reliability for clinical and forensic applications. Therefore, ChatGPT-based estimations should be interpreted with caution until future versions achieve more consistent and reproducible performance through population-specific training, model optimization, and multicenter validation.