Purpose <p>To compare deep learning models of different architecture for automated lumbar spinal stenosis classification on MRI and benchmark their performance against radiologists and orthopedists.</p> Methods <p>Lumbar spine MRI studies from Sep-2015 to Sep-2019 were retrospectively obtained. Exclusion criteria included previous spinal instrumentation, suboptimal image quality, post-gadolinium studies, and severe scoliosis. Axial T2-weighted and sagittal T1-weighted images were used. Studies were split into training/validation and test sets. An external test set of 100 studies was used. Training data were labelled by 4 radiologists using predefined gradings. Two models, CNN-based and transformer-based, were developed. Consensus labelling by two expert spine radiologists served as the reference standard. Test sets were labelled by 8 participants (2 general radiologists, 2 radiologists-in-training, 2 orthopedists, 2 orthopedists-in-training). Detection recall (%), interrater agreement (Gwet κ), sensitivity, and specificity were evaluated.</p> Results <p>564 MRI lumbar spines were included (mean age = 52 ± 19[SD]; 302 women), with 464(82%) and 100(18%) for training/validation and internal testing, respectively. Both models showed high recall for all regions of interest (&gt; 94%), similar to participants. Dichotomous classification (normal/mild vs. moderate/severe) by the CNN model, transformer model, and participants showed respective kappas for central canal 0.99/0.99/0.97–0.98, lateral recesses 0.98/0.94/0.81–0.94, and neural foramina 0.98/0.95/0.91–0.95 on internal testing (<i>p</i> &lt; 0.001); for central canal 0.99/0.97/0.92–0.97, lateral recess 0.97/0.90/0.61–0.91, and neural foramina 0.99/0.94/0.87–0.93 on external testing (<i>p</i> &lt; 0.001).</p> Conclusion <p>The CNN model showed superior performance, and the transformer model showed similar to superior performance compared to clinicians for classifying lumbar spinal stenosis. These models could assist clinicians in report generation, surgical planning and education.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep learning models for lumbar spinal stenosis on MRI: model comparison and clinical benchmarking

  • You Jun Lee,
  • Changshuo Liu,
  • Yong Han Ting,
  • Andrew Makmur,
  • Weizhong Jonathan Sng,
  • Amanda J. L. Cheng,
  • Jiong Hao Tan,
  • Alex Quok An Teo,
  • Zongchen Li,
  • Alvin Hong Zhi Ng,
  • Aric Lee,
  • Chongyan Wang,
  • Xinyi Lim,
  • Qai Ven Yap,
  • Joey Chan Yiing Beh,
  • Shuxun Lin,
  • Naresh Kumar,
  • Beng Chin Ooi,
  • James Thomas Patrick Decourcy Hallinan

摘要

Purpose

To compare deep learning models of different architecture for automated lumbar spinal stenosis classification on MRI and benchmark their performance against radiologists and orthopedists.

Methods

Lumbar spine MRI studies from Sep-2015 to Sep-2019 were retrospectively obtained. Exclusion criteria included previous spinal instrumentation, suboptimal image quality, post-gadolinium studies, and severe scoliosis. Axial T2-weighted and sagittal T1-weighted images were used. Studies were split into training/validation and test sets. An external test set of 100 studies was used. Training data were labelled by 4 radiologists using predefined gradings. Two models, CNN-based and transformer-based, were developed. Consensus labelling by two expert spine radiologists served as the reference standard. Test sets were labelled by 8 participants (2 general radiologists, 2 radiologists-in-training, 2 orthopedists, 2 orthopedists-in-training). Detection recall (%), interrater agreement (Gwet κ), sensitivity, and specificity were evaluated.

Results

564 MRI lumbar spines were included (mean age = 52 ± 19[SD]; 302 women), with 464(82%) and 100(18%) for training/validation and internal testing, respectively. Both models showed high recall for all regions of interest (> 94%), similar to participants. Dichotomous classification (normal/mild vs. moderate/severe) by the CNN model, transformer model, and participants showed respective kappas for central canal 0.99/0.99/0.97–0.98, lateral recesses 0.98/0.94/0.81–0.94, and neural foramina 0.98/0.95/0.91–0.95 on internal testing (p < 0.001); for central canal 0.99/0.97/0.92–0.97, lateral recess 0.97/0.90/0.61–0.91, and neural foramina 0.99/0.94/0.87–0.93 on external testing (p < 0.001).

Conclusion

The CNN model showed superior performance, and the transformer model showed similar to superior performance compared to clinicians for classifying lumbar spinal stenosis. These models could assist clinicians in report generation, surgical planning and education.