错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lightweight Rendezvous Model Based on Pruning and Knowledge Distillation for Action Triplet Recognition in Laparoscopic Surgery Videos

  • Manh-Hung Ha,
  • Kim Dinh Thai,
  • Dang Cong Vinh

摘要

This study presents a novel approach using Deep Neural Networks (DNNs) based on Rendezvous (RDV) model to accurately identify action triplets in laparoscopic surgical videos. Our DNN architecture consists of sequentially arranged neuron layers, including an input layer containing neurons representing input data, hidden layers containing multiple neurons and corresponding weights to learn features, and an output layer representing prediction results. Particular, the RDV model directly identifies the triplets from surgical videos by combining Attention mechanisms at two different levels: Class Activation Guided Attention Mechanism (CAGAM) and Multi-Head of Mixed Attention (MHMA). Our model focuses on developing an efficient attention mechanism to enhance the identification and classification surgical action triplets in endoscopic videos. Additionally, to enhance performance and optimize the RDV model, we use techniques such as Pruning and Knowledge Distillation, achieving significantly improved prediction time compared to the original model. In the experiment, the results show that our proposed model performs on the CholecT50 dataset (which includes 50 laparoscopic cholecystectomy videos, with each frame labeled with triplets from 100 different classes) with average accuracies of 0.8397, 0.4633, 0.3024, and 0.2095 for tool classification, action, target, and triplets, respectively.