错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Harnessing Temporal Information for Precise Frame-Level Predictions in Endoscopy Videos

  • Pooya Mobadersany,
  • Chaitanya Parmar,
  • Pablo F. Damasceno,
  • Shreyas Fadnavis,
  • Krishna Chaitanya,
  • Shilong Li,
  • Evan Schwab,
  • Jaclyn Xiao,
  • Lindsey Surace,
  • Tommaso Mansi,
  • Gabriela Oana Cula,
  • Louis R. Ghanem,
  • Kristopher Standish

摘要

Camera localization in endoscopy videos plays a fundamental role in enabling precise diagnosis and effective treatment planning for patients with Inflammatory Bowel Disease (IBD). Precise frame-level classification, however, depends on long-range temporal dynamics, ranging from hundreds to tens of thousands of frames per video, challenging current neural network approaches. To address this, we propose EndoFormer, a frame-level classification model that leverages long-range temporal information for anatomic segment classification in gastrointestinal endoscopy videos. EndoFormer combines a Foundation Model block, judicious video-level augmentations, and a Transformer classifier for frame-level classification while maintaining a small memory footprint. Experiments on 4160 endoscopy videos from four clinical trials and over 61 million frames demonstrate that EndoFormer has an AUC = 0.929, significantly improving state-of-the-art models for anatomic segment classification. These results highlight the potential for adopting EndoFormer in endoscopy video analysis applications that require long-range temporal dynamics for precise frame-level predictions.