<p>Better scene understanding of historical districts is critically significant for heritage protection and urban development. However, current methods, manual or automatic, fail to effectively capture street-view entities and their relationships based on historical street-view images due to panorama imaging approaches and changing brightness conditions. To address this issue, this study introduces a Historical Street-view Interpretation Network (HSVI-Net) based on panoramic images. HSVI-Net innovatively incorporates enhanced channel, space, coordinate, and multi-scale attention mechanisms to model semantic information within images, thereby directly predicting the triplets for scene understanding. The experimental results demonstrated HSVI-Net’s superior performance on the historical street-view dataset, improving recall, mean recall, AP50, and prediction accuracy by up to 5.05%, 1.22%, 11.30%, and 22.86% respectively, and facilitating accurate understanding of complex historical street-view with a lightweight advantage. These advancements benefit for digital representation and modelling analysis of historical districts for the participants of protection, revitalization, and utilization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HSVI-Net: a deep learning network for scene understanding based on panoramic street-view images within historical districts

  • Huajian Zhang,
  • Jie Jiang,
  • Xian Guo

摘要

Better scene understanding of historical districts is critically significant for heritage protection and urban development. However, current methods, manual or automatic, fail to effectively capture street-view entities and their relationships based on historical street-view images due to panorama imaging approaches and changing brightness conditions. To address this issue, this study introduces a Historical Street-view Interpretation Network (HSVI-Net) based on panoramic images. HSVI-Net innovatively incorporates enhanced channel, space, coordinate, and multi-scale attention mechanisms to model semantic information within images, thereby directly predicting the triplets for scene understanding. The experimental results demonstrated HSVI-Net’s superior performance on the historical street-view dataset, improving recall, mean recall, AP50, and prediction accuracy by up to 5.05%, 1.22%, 11.30%, and 22.86% respectively, and facilitating accurate understanding of complex historical street-view with a lightweight advantage. These advancements benefit for digital representation and modelling analysis of historical districts for the participants of protection, revitalization, and utilization.