This study aims to enhance the detection of small objects within intricate visual environments by harnessing the latest deep-learning advancements. We propose a novel method that features a dual-stream self-attention mechanism integrated within a multi-head framework, along with an innovative output reweighting technique to further refine detection accuracy. The core of our approach is specifically designed to address the difficulties posed by small objects that often overlap multiple tokens in feature maps, a common limitation of traditional detection models. By dynamically adjusting the attention scale across various heads, our method facilitates detailed feature capture at multiple levels of granularity, significantly improving the model’s capability to detect and describe small objects. Additionally, we introduce a softmax-based reweighting function that selectively emphasizes crucial features for object recognition, reducing noise and irrelevant information. Our model, named SSD-MSDSSA-ORT, not only exceeds the accuracy of existing state-of-the-art solutions but also showcases superior processing efficiency and scalability. These contributions advance the theoretical understanding of attention mechanisms in deep neural networks and offer practical enhancements for real-world applications requiring precise small object detection.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Modified Attention Block for Detecting Cars at a Great Distance

  • A. Samarin,
  • A. Savelev,
  • A. Toropov,
  • A. Dzestelova,
  • A. Motyko,
  • E. Kotenko,
  • E. Mikhailova,
  • V. Malykh

摘要

This study aims to enhance the detection of small objects within intricate visual environments by harnessing the latest deep-learning advancements. We propose a novel method that features a dual-stream self-attention mechanism integrated within a multi-head framework, along with an innovative output reweighting technique to further refine detection accuracy. The core of our approach is specifically designed to address the difficulties posed by small objects that often overlap multiple tokens in feature maps, a common limitation of traditional detection models. By dynamically adjusting the attention scale across various heads, our method facilitates detailed feature capture at multiple levels of granularity, significantly improving the model’s capability to detect and describe small objects. Additionally, we introduce a softmax-based reweighting function that selectively emphasizes crucial features for object recognition, reducing noise and irrelevant information. Our model, named SSD-MSDSSA-ORT, not only exceeds the accuracy of existing state-of-the-art solutions but also showcases superior processing efficiency and scalability. These contributions advance the theoretical understanding of attention mechanisms in deep neural networks and offer practical enhancements for real-world applications requiring precise small object detection.