Purpose <p>To assess the performance of state-of-the-art large language models in classifying vertebral metastasis stability using the Spinal Instability Neoplastic Score (SINS) compared to human experts, and to evaluate the impact of task-specific refinement including in-context learning on their performance.</p> Material and methods <p>This retrospective study analyzed 100 synthetic CT and MRI reports encompassing a broad range of SINS scores. Four human experts (two radiologists and two neurosurgeons) and four large language models (Mistral, Claude, GPT-4 turbo, and GPT-4o) evaluated the reports. Large language models were tested in both generic form and with task-specific refinement. Performance was assessed based on correct SINS category assignment and attributed SINS points.</p> Results <p>Human experts demonstrated high median performance in SINS classification (98.5% correct) and points calculation (92% correct), with a median point offset of 0 [0–0]. Generic large language models performed poorly with 26–63% correct category and 4–15% correct SINS points allocation. In-context learning significantly improved chatbot performance to near-human levels (96–98/100 correct for classification, 86–95/100 for scoring, no significant difference to human experts). Refined large language models performed 71–85% better in SINS points allocation.</p> Conclusion <p>In-context learning enables state-of-the-art large language models to perform at near-human expert levels in SINS classification, offering potential for automating vertebral metastasis stability assessment. The poor performance of generic large language models highlights the importance of task-specific refinement in medical applications of artificial intelligence.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

In-context learning enables large language models to achieve human-level performance in spinal instability neoplastic score classification from synthetic CT and MRI reports

  • Maximilian F. Russe,
  • Marco Reisert,
  • Anna Fink,
  • Marc Hohenhaus,
  • Julia M. Nakagawa,
  • Caroline Wilpert,
  • Carl P. Simon,
  • Elmar Kotter,
  • Horst Urbach,
  • Alexander Rau

摘要

Purpose

To assess the performance of state-of-the-art large language models in classifying vertebral metastasis stability using the Spinal Instability Neoplastic Score (SINS) compared to human experts, and to evaluate the impact of task-specific refinement including in-context learning on their performance.

Material and methods

This retrospective study analyzed 100 synthetic CT and MRI reports encompassing a broad range of SINS scores. Four human experts (two radiologists and two neurosurgeons) and four large language models (Mistral, Claude, GPT-4 turbo, and GPT-4o) evaluated the reports. Large language models were tested in both generic form and with task-specific refinement. Performance was assessed based on correct SINS category assignment and attributed SINS points.

Results

Human experts demonstrated high median performance in SINS classification (98.5% correct) and points calculation (92% correct), with a median point offset of 0 [0–0]. Generic large language models performed poorly with 26–63% correct category and 4–15% correct SINS points allocation. In-context learning significantly improved chatbot performance to near-human levels (96–98/100 correct for classification, 86–95/100 for scoring, no significant difference to human experts). Refined large language models performed 71–85% better in SINS points allocation.

Conclusion

In-context learning enables state-of-the-art large language models to perform at near-human expert levels in SINS classification, offering potential for automating vertebral metastasis stability assessment. The poor performance of generic large language models highlights the importance of task-specific refinement in medical applications of artificial intelligence.