HSVI-Net: a deep learning network for scene understanding based on panoramic street-view images within historical districts
摘要
Better scene understanding of historical districts is critically significant for heritage protection and urban development. However, current methods, manual or automatic, fail to effectively capture street-view entities and their relationships based on historical street-view images due to panorama imaging approaches and changing brightness conditions. To address this issue, this study introduces a Historical Street-view Interpretation Network (HSVI-Net) based on panoramic images. HSVI-Net innovatively incorporates enhanced channel, space, coordinate, and multi-scale attention mechanisms to model semantic information within images, thereby directly predicting the triplets for scene understanding. The experimental results demonstrated HSVI-Net’s superior performance on the historical street-view dataset, improving recall, mean recall, AP50, and prediction accuracy by up to 5.05%, 1.22%, 11.30%, and 22.86% respectively, and facilitating accurate understanding of complex historical street-view with a lightweight advantage. These advancements benefit for digital representation and modelling analysis of historical districts for the participants of protection, revitalization, and utilization.