Introduction <p>Cut-out is the most consequential mechanical complication after proximal femoral nailing and requires prompt recognition. General-purpose artificial intelligence (AI) models with image-interpretation capability are increasingly accessible, yet their diagnostic performance for this task is unknown. We evaluated ChatGPT’s accuracy for predicting cut-out following proximal femoral nailing and compared TFNA and Gamma nail subgroups.</p> Materials and methods <p>In this retrospective predictive-accuracy study, 989 patients (683 TFNA, 306 Gamma nail; mean age 83.8 years; cut-out prevalence 2.5%) were analysed. For each case, three baseline radiographs — the injury anteroposterior and lateral views and the immediate postoperative control radiograph, all obtained before any complication could be radiographically present — together with the full clinical data set (excluding complication-related information) were submitted to ChatGPT, which returned a binary prediction (yes/no) and probability estimate (0–100%). The follow-up radiographs on which cut-out becomes apparent were not provided to the model. Ground truth was established on subsequent follow-up radiographs by a departmental review board requiring agreement of two senior surgeons. Accuracy metrics were calculated with Wilson 95% confidence intervals (CI); exact McNemar’s and Mann-Whitney U tests assessed directional bias and calibration. Reporting followed STARD guidelines.</p> Results <p>ChatGPT showed a sensitivity of 68.0% (95% CI 48.4–82.8%), specificity of 62.7% (59.6–65.7%), PPV of 4.5% (2.8–7.1%), NPV of 98.7% (97.4–99.3%), and AUC of 0.694 (0.580–0.790). Performance was higher for TFNA (sensitivity 73.7%, AUC 0.730) than for Gamma nail (sensitivity 50.0%, AUC 0.606). The model over-predicted cut-out, with 360 false positives against 8 false negatives (45:1; McNemar’s <i>p</i> &lt; 0.001, Holm-corrected). Calibration was inverted: the median predicted probability was lower for correct (9%) than for incorrect (40%) classifications (Mann-Whitney <i>p</i> &lt; 0.001).</p> Conclusions <p>ChatGPT demonstrated limited, clinically insufficient accuracy for predicting cut-out following proximal femoral nailing, with a prohibitive false-positive burden and inverted calibration. Subgroup point estimates were lower for Gamma nail cases, but this exploratory difference did not reach statistical significance. General-purpose AI models are not currently suitable as a substitute for, or adjunct to, surgeon-led surveillance of this complication. Task-specific training, external validation, and defined clinical boundaries are required before such tools can be considered for fracture follow-up pathways.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predictive accuracy of a general-purpose artificial intelligence model for cut-out following proximal femoral nailing

  • Guy Ben Arie,
  • Yaniv Warschawski,
  • Nissan Amzallag,
  • Hadar Gan-Or,
  • Tomer Ben-Tov,
  • Amal Khoury,
  • Idit Horev,
  • Nadav Graif,
  • Guy Morag

摘要

Introduction

Cut-out is the most consequential mechanical complication after proximal femoral nailing and requires prompt recognition. General-purpose artificial intelligence (AI) models with image-interpretation capability are increasingly accessible, yet their diagnostic performance for this task is unknown. We evaluated ChatGPT’s accuracy for predicting cut-out following proximal femoral nailing and compared TFNA and Gamma nail subgroups.

Materials and methods

In this retrospective predictive-accuracy study, 989 patients (683 TFNA, 306 Gamma nail; mean age 83.8 years; cut-out prevalence 2.5%) were analysed. For each case, three baseline radiographs — the injury anteroposterior and lateral views and the immediate postoperative control radiograph, all obtained before any complication could be radiographically present — together with the full clinical data set (excluding complication-related information) were submitted to ChatGPT, which returned a binary prediction (yes/no) and probability estimate (0–100%). The follow-up radiographs on which cut-out becomes apparent were not provided to the model. Ground truth was established on subsequent follow-up radiographs by a departmental review board requiring agreement of two senior surgeons. Accuracy metrics were calculated with Wilson 95% confidence intervals (CI); exact McNemar’s and Mann-Whitney U tests assessed directional bias and calibration. Reporting followed STARD guidelines.

Results

ChatGPT showed a sensitivity of 68.0% (95% CI 48.4–82.8%), specificity of 62.7% (59.6–65.7%), PPV of 4.5% (2.8–7.1%), NPV of 98.7% (97.4–99.3%), and AUC of 0.694 (0.580–0.790). Performance was higher for TFNA (sensitivity 73.7%, AUC 0.730) than for Gamma nail (sensitivity 50.0%, AUC 0.606). The model over-predicted cut-out, with 360 false positives against 8 false negatives (45:1; McNemar’s p < 0.001, Holm-corrected). Calibration was inverted: the median predicted probability was lower for correct (9%) than for incorrect (40%) classifications (Mann-Whitney p < 0.001).

Conclusions

ChatGPT demonstrated limited, clinically insufficient accuracy for predicting cut-out following proximal femoral nailing, with a prohibitive false-positive burden and inverted calibration. Subgroup point estimates were lower for Gamma nail cases, but this exploratory difference did not reach statistical significance. General-purpose AI models are not currently suitable as a substitute for, or adjunct to, surgeon-led surveillance of this complication. Task-specific training, external validation, and defined clinical boundaries are required before such tools can be considered for fracture follow-up pathways.