<b>Purpose</b> <p>While significant progress has been made in skill assessment for minimally invasive procedures, objective evaluation methods for open surgery remain limited. This paper presents a deep learning framework for assessing technical surgical skills using egocentric video data from open surgery training.</p> <b>Methods</b> <p>Our dataset includes 201 videos and corresponding hand kinematics data from three fundamental training task—knot tying (KT), continuous suturing (CS), and interrupted suturing (IS)—performed by 20 participants. Each video was annotated by two experts using a modified OSATS scale (KT: five criteria, total score range: 5–25; CS/IS: seven criteria, total score range: 7–35). We evaluate three temporal architectures (LSTM, TCN, and Transformer), each using ResNet50 as the backbone for spatial feature extraction, and assess them under various training strategies: single-task learning, feature concatenation, pretraining, and multi-task learning with integrated kinematic data. Performance metrics included mean absolute error (MAE) and Spearman correlation coefficient (<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\rho \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ρ</mi> </math></EquationSource> </InlineEquation>), both with respect to total score prediction.</p> <b>Results</b> <p>The Transformer-based models consistently outperformed LSTM and TCN across all tasks. The multi-task Transformer incorporating prediction of task completion time (<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\text {Transf-MT}_{\text {T+S}}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mtext>Transf-MT</mtext> <mtext>T+S</mtext> </msub> </math></EquationSource> </InlineEquation>) achieved the lowest MAE (KT: 1.92, CS: 2.81, and IS: 2.89) and <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\rho \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ρ</mi> </math></EquationSource> </InlineEquation> = 0.84-<InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(-\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>-</mo> </math></EquationSource> </InlineEquation>0.90. It also demonstrated promising capabilities for early skill assessment by predicting the total score from partial observations—particularly for simpler tasks. Additionally, we show that models trained on consensus expert ratings outperform those trained on individual annotations, highlighting the value of multi-rater ground truth.</p> <b>Conclusion</b> <p>This research provides a foundation for objective, automated assessment of open surgical skills, with potential to improve the efficiency and standardization of surgical training.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Egocentric video analysis for automated assessment of open surgical skills via deep learning

  • Athanasios Gazis,
  • Dimitrios Schizas,
  • Stylianos Kykalos,
  • Pantelis Karaiskos,
  • Constantinos Loukas

摘要

Purpose

While significant progress has been made in skill assessment for minimally invasive procedures, objective evaluation methods for open surgery remain limited. This paper presents a deep learning framework for assessing technical surgical skills using egocentric video data from open surgery training.

Methods

Our dataset includes 201 videos and corresponding hand kinematics data from three fundamental training task—knot tying (KT), continuous suturing (CS), and interrupted suturing (IS)—performed by 20 participants. Each video was annotated by two experts using a modified OSATS scale (KT: five criteria, total score range: 5–25; CS/IS: seven criteria, total score range: 7–35). We evaluate three temporal architectures (LSTM, TCN, and Transformer), each using ResNet50 as the backbone for spatial feature extraction, and assess them under various training strategies: single-task learning, feature concatenation, pretraining, and multi-task learning with integrated kinematic data. Performance metrics included mean absolute error (MAE) and Spearman correlation coefficient ( \(\rho \) ρ ), both with respect to total score prediction.

Results

The Transformer-based models consistently outperformed LSTM and TCN across all tasks. The multi-task Transformer incorporating prediction of task completion time ( \(\text {Transf-MT}_{\text {T+S}}\) Transf-MT T+S ) achieved the lowest MAE (KT: 1.92, CS: 2.81, and IS: 2.89) and \(\rho \) ρ = 0.84- \(-\) - 0.90. It also demonstrated promising capabilities for early skill assessment by predicting the total score from partial observations—particularly for simpler tasks. Additionally, we show that models trained on consensus expert ratings outperform those trained on individual annotations, highlighting the value of multi-rater ground truth.

Conclusion

This research provides a foundation for objective, automated assessment of open surgical skills, with potential to improve the efficiency and standardization of surgical training.