<p>In underwater environments, transparent organisms with low visibility and minimal visual features, lacking distinctive shadows or silhouettes, can blend seamlessly into their surroundings. Existing deep learning methods for detecting such organisms have shown unsatisfactory performance. This study proposes a multimodal fusion network, UTNet, which combines event-based and red-green-blue (RGB)-based vision for the underwater transparent camouflaged organism detection task. UTNet introduces a two-stage enhanced representation aggregation module comprising a multi-feature aggregation component (MFAC) and a deep fusion component (DFC) to facilitate the synergy between frame-based and event-based vision. First, MFAC aggregates the high dynamic range features from events with the static details from RGB images. Then, the edge information from the edge clue search module is used to guide the fusion process, reducing background interference. Next, DFC further extracts depth information from the MFAC output using five parallel branches. Additionally, a submanifold sparse convolution-modified ResNet50 backbone network is employed to extract features from event frames, preserving event sparsity and improving computational efficiency. Extensive experiments on our custom underwater transparent organism dataset, captured using the DAVIS346 event camera, demonstrate the effectiveness of UTNet. The results show that UTNet achieves 75.2% accuracy and 37.8 frames per second, providing the best trade-off between speed and accuracy compared to other detectors.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

UTNet: event-RGB multimodal fusion model for underwater transparent organism detection

  • Fengyue Guo,
  • Peng Ren,
  • Cai Luo

摘要

In underwater environments, transparent organisms with low visibility and minimal visual features, lacking distinctive shadows or silhouettes, can blend seamlessly into their surroundings. Existing deep learning methods for detecting such organisms have shown unsatisfactory performance. This study proposes a multimodal fusion network, UTNet, which combines event-based and red-green-blue (RGB)-based vision for the underwater transparent camouflaged organism detection task. UTNet introduces a two-stage enhanced representation aggregation module comprising a multi-feature aggregation component (MFAC) and a deep fusion component (DFC) to facilitate the synergy between frame-based and event-based vision. First, MFAC aggregates the high dynamic range features from events with the static details from RGB images. Then, the edge information from the edge clue search module is used to guide the fusion process, reducing background interference. Next, DFC further extracts depth information from the MFAC output using five parallel branches. Additionally, a submanifold sparse convolution-modified ResNet50 backbone network is employed to extract features from event frames, preserving event sparsity and improving computational efficiency. Extensive experiments on our custom underwater transparent organism dataset, captured using the DAVIS346 event camera, demonstrate the effectiveness of UTNet. The results show that UTNet achieves 75.2% accuracy and 37.8 frames per second, providing the best trade-off between speed and accuracy compared to other detectors.