In the field of endoscopic video surgical workflow analysis, surgical action triplet recognition is a comprehensive fine-grained surgical activity recognition task. It establishes data associations among instruments, actions, and targets, providing a standardized description of the interactions between instruments and targets. Due to the lack of spatial annotations, existing methods mostly use weakly supervised learning to locate instruments, which prevents the models from fully utilizing the spatial information in surgical videos, thus reducing the accuracy of triplet recognition. To address this challenge, we ingeniously applied the medical Foundation Model MedSAM [14] to generate spatial annotations of the surgical instruments, providing technical support for precise instrument localization. At the same time, we introduced a supervision learning method based on pseudo-labels, incorporating an instrument segmentation task into the triplet recognition task, using the pseudo-labels generated by the medical Foundation Model as the supervisory signal for the segmentation module. This approach enhances the model’s accuracy in locating surgical instruments, thereby improving the overall precision of triplet recognition. We evaluated our model on the CholecT50 dataset, and compared with the baseline model that uses weakly supervised methods for instrument localization, there was a significant improvement in various indicators of action triplet recognition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Surgical Action Triplet Recognition Assisted by Foundation Models-Based Instrument Localization

  • Junyan Liu,
  • Peng Qiao,
  • Yong Dou,
  • Sidun Liu,
  • Lu Shen,
  • Xi Wang,
  • Wenyu Li

摘要

In the field of endoscopic video surgical workflow analysis, surgical action triplet recognition is a comprehensive fine-grained surgical activity recognition task. It establishes data associations among instruments, actions, and targets, providing a standardized description of the interactions between instruments and targets. Due to the lack of spatial annotations, existing methods mostly use weakly supervised learning to locate instruments, which prevents the models from fully utilizing the spatial information in surgical videos, thus reducing the accuracy of triplet recognition. To address this challenge, we ingeniously applied the medical Foundation Model MedSAM [14] to generate spatial annotations of the surgical instruments, providing technical support for precise instrument localization. At the same time, we introduced a supervision learning method based on pseudo-labels, incorporating an instrument segmentation task into the triplet recognition task, using the pseudo-labels generated by the medical Foundation Model as the supervisory signal for the segmentation module. This approach enhances the model’s accuracy in locating surgical instruments, thereby improving the overall precision of triplet recognition. We evaluated our model on the CholecT50 dataset, and compared with the baseline model that uses weakly supervised methods for instrument localization, there was a significant improvement in various indicators of action triplet recognition.