Efficient communication is imperative to ensure resilient and real-time multi-agent collaborative perception that presents as an emerging application in autonomous driving. This paper proposes the AttenS, a conceptually simple yet effective intermediate collaborative perception framework, striking the balance between performances and communication volumes. Differing from previous works that focusing on channel-wise compression or pixel-wise spatial filtering, AttenS leverages patch-wise spatial filtering which is more appropriate attributed to local semantic information and inherent localization error resilience of patches compared to isolated pixels. Concretely, we design a Vision Transformer-like method to attentively select the perceptually critical areas, which can support effective spatial compression. Moreover, a local attention-based module is constructed to refine the spatial inconsistency introduced by the spatially discontinuous feature aggregation. Eventually, AttenS is validated on V2XSet dataset, a simulated Vehicle-to-Everything dataset. During experiments, AttenS presents top-rank performances and communication efficacy compared to previous works, thereby its effectiveness and superiority are demonstrated.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AttenS: An Attentive Selection Method for Communication-Efficient Cooperative Perception

  • Yijie Chen,
  • Yuzhe Ji,
  • Chenyu Sun,
  • Rui Ding,
  • Meng Yang,
  • Xinhu Zheng

摘要

Efficient communication is imperative to ensure resilient and real-time multi-agent collaborative perception that presents as an emerging application in autonomous driving. This paper proposes the AttenS, a conceptually simple yet effective intermediate collaborative perception framework, striking the balance between performances and communication volumes. Differing from previous works that focusing on channel-wise compression or pixel-wise spatial filtering, AttenS leverages patch-wise spatial filtering which is more appropriate attributed to local semantic information and inherent localization error resilience of patches compared to isolated pixels. Concretely, we design a Vision Transformer-like method to attentively select the perceptually critical areas, which can support effective spatial compression. Moreover, a local attention-based module is constructed to refine the spatial inconsistency introduced by the spatially discontinuous feature aggregation. Eventually, AttenS is validated on V2XSet dataset, a simulated Vehicle-to-Everything dataset. During experiments, AttenS presents top-rank performances and communication efficacy compared to previous works, thereby its effectiveness and superiority are demonstrated.