<p>Traditional DETR performs poorly in small target detection tasks, while frequency domain features can supplement spatial domain information and enrich image expression, thus significantly improving the accuracy of small target detection. Therefore, this paper introduces a DETR model named Signal-DETR, which integrates the Fourier transform with multi-head attention mechanisms. The model uses the Fourier transform to extract signal information for each query, particularly the phase and magnitude, and incorporates this auxiliary information into the query for subsequent attention operations. The phase component preserves the geometric structure, spatial arrangement, and intricate details of the features, thereby ensuring the continuity of lane detection. Meanwhile, the magnitude component provides insights into the intensity distribution of the features, aiding the model in capturing global characteristics and precise lane positions. By seamlessly merging the phase and magnitude with multi-head attention, the model enhances its representation of both global and low-frequency features. This fusion bolsters its capacity to perceive multi-scale features and increases robustness in handling geometric transformations and intricate scenarios, ultimately achieving more accurate lane detection. Experimental results demonstrate that Signal-DETR, which utilizes a ResNet50 backbone, attains an F1 score of 78.44% on the CULane dataset. This performance marks a significant advancement compared to existing Transformers and CNN-based detectors.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing lane line detection transformer with frequency-assisted attention

  • Xuandong Zhao,
  • Wei Wu

摘要

Traditional DETR performs poorly in small target detection tasks, while frequency domain features can supplement spatial domain information and enrich image expression, thus significantly improving the accuracy of small target detection. Therefore, this paper introduces a DETR model named Signal-DETR, which integrates the Fourier transform with multi-head attention mechanisms. The model uses the Fourier transform to extract signal information for each query, particularly the phase and magnitude, and incorporates this auxiliary information into the query for subsequent attention operations. The phase component preserves the geometric structure, spatial arrangement, and intricate details of the features, thereby ensuring the continuity of lane detection. Meanwhile, the magnitude component provides insights into the intensity distribution of the features, aiding the model in capturing global characteristics and precise lane positions. By seamlessly merging the phase and magnitude with multi-head attention, the model enhances its representation of both global and low-frequency features. This fusion bolsters its capacity to perceive multi-scale features and increases robustness in handling geometric transformations and intricate scenarios, ultimately achieving more accurate lane detection. Experimental results demonstrate that Signal-DETR, which utilizes a ResNet50 backbone, attains an F1 score of 78.44% on the CULane dataset. This performance marks a significant advancement compared to existing Transformers and CNN-based detectors.