Evaluating large language models for surgical chart review of second stage implant-based breast reconstruction: a comparative analysis of manual review, GPT-3.5 Turbo, and GPT-4 Turbo
摘要
Manual chart review is labor-intensive and error-prone. Integrating Large Language Models (LLMs), like GPT, may reduce errors and accelerate the scalability of plastic surgery studies. This study compares the performance of GPT-3.5 Turbo and GPT-4 Turbo to manual review for second stage implant-based breast reconstructions requiring capsule revision.
MethodsThis retrospective cohort study identified 101 s stage breast reconstruction surgeries at Stanford Hospital (2018–2021) using CPT codes. Chart reviews collected patient demographics and operative details. Capsule procedures included capsulectomy, capsulorrhaphy, and ‘extensive’ capsulotomy, defined as any capsulotomy that changed the breast pocket position beyond entry into the capsule. GPT-3.5 Turbo and GPT-4 Turbo filtered operative reports and identified revisions. Performance metrics were averaged over 10 runs for each LLM.
ResultsThe GPT-3.5 Turbo model performed well in determining whether a second stage breast reconstruction occurred, incorrectly categorizing just one operative report. The model correctly predicted 85.2% of capsulectomies, 76.9% of capsulorrhaphies and 78.1% of extensive capsulotomies. The overall success rate was 80.1%, with a recall of 0.76, precision of 0.71, and F-score of 0.72. Re-running with GPT-4 Turbo considerably improved performance, correctly identifying 86.9% of capsulectomies, 88.3% of capsulorrhaphies, and 85.3% of extensive capsulotomies, with an overall success rate of 86.8%, recall of 0.94, precision of 0.69, and F-score of 0.78.
ConclusionsWhile GPT-4 Turbo outperformed GPT-3.5 Turbo, both models exhibit limitations in precisely matching manual chart review. These findings emphasize the potential as first pass tools, but highlight the need for cautious interpretation and complementary manual review in ensuring accurate data collection.
Level of EvidenceNot gradable.