Multi-view methods offer an effective approach to address 3D object classification, and there is a growing interest in handling more practical scenarios, such as fine-grained distinctions and arbitrary viewpoints. Fine-grained object distinctions often appearance around local parts, inspiring various part-based methods. However, the parts generated by these methods are typically unordered, significantly impacting the aggregation operation across different views and consequently diminishing overall performance. To address this issue, we propose a Part2Part Hierarchical Attention, facilitating information exchange and feature enhancement among parts in different viewpoints through attention within and across views. Subsequently, the Part2Part Attention Map generated during the attention process is utilized to measure the distance between multi-view parts, aiding in their alignment. We also employ multi-scale feature fusion to enhance the quality of parts generated by weakly supervised learning. Experimental results indicate that, under the same settings, our approach achieves state-of-the-art performance on the FG3D and MVP-N datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention4Align: Align Multi-view Parts Via Part2Part Hierarchical Attention Map for Fine-Grained 3D Object Classification

  • Runchu Zhang,
  • Jiahe Yue,
  • Zhe Zhang,
  • Jie Ma

摘要

Multi-view methods offer an effective approach to address 3D object classification, and there is a growing interest in handling more practical scenarios, such as fine-grained distinctions and arbitrary viewpoints. Fine-grained object distinctions often appearance around local parts, inspiring various part-based methods. However, the parts generated by these methods are typically unordered, significantly impacting the aggregation operation across different views and consequently diminishing overall performance. To address this issue, we propose a Part2Part Hierarchical Attention, facilitating information exchange and feature enhancement among parts in different viewpoints through attention within and across views. Subsequently, the Part2Part Attention Map generated during the attention process is utilized to measure the distance between multi-view parts, aiding in their alignment. We also employ multi-scale feature fusion to enhance the quality of parts generated by weakly supervised learning. Experimental results indicate that, under the same settings, our approach achieves state-of-the-art performance on the FG3D and MVP-N datasets.