TransSG: A Spatial-Temporal Transformer for Surgical Gesture Recognition
摘要
The goal of this research is to identify surgical movements using RGB videos. Prior methods for recognizing surgical gestures mainly concentrate on robot-assisted surgery, but less on surgical gesture training. In this paper, we introduce a novel surgical gesture dataset that comprises the human surgical gestures of experienced surgeons captured from multiple viewpoints. Moreover, we propose a spatial-temporal transformer-based surgical gesture recognition model called TransSG to identify the user’s surgical gestures better. Our framework comprises three key elements: (1) the importance-weight layer that learns the significance of segmented patches before passing them to the spatial encoder, (2) the aggregated temporal feature extractor that facilitates computationally efficient processing, and (3) a novel loss function that integrates a range of effective loss functions to enhance recognition quality. Our experiments demonstrate that TransSG attains excellent Top-1 accuracy (83.6%), surpassing existing methods on the proposed surgical gesture dataset.