Spatio-temporal Context Fusion Network for Endoscopic Ultrasound Video Segmentation
摘要
The accurate and automatic segmentation of endoscopic ultrasound (EUS) videos is crucial for assessing the progression of submucosal tumors (SMTs) in the digestive tract and developing treatment plans. However, the dynamic nature of these videos and the rich temporal information that they contain render traditional static image processing techniques insufficient for fully extracting potential diagnostic information. To address this issue, a spatio-temporal aggregation network is proposed for the precise segmentation of EUS videos. This network comprises a spatio-temporal attention module (STAM) and a spatio-temporal context fusion module (STFM). The STAM extracts spatio-temporal features using attention mechanisms in both the temporal and spatial dimensions. The STFM synchronously integrates the spatial context from static images with the temporal context across different frames to fully explore the potential diagnostic information. The experimental results demonstrate that this method significantly improves the semantic segmentation performance of SMT videos, such as those of gastrointestinal stromal tumors and leiomyomas.