ViT-HVE: a vision transformer-based framework for recognition and weighted evaluation of cultural heritage values
摘要
Cultural heritage value is essential for heritage protection and utilization. However, existing evaluation methods remain subjective and labor-intensive. To address these challenges, we propose the Vision Transformer-based Heritage Value Evaluation (ViT-HVE) model, the first deep learning-based model for heritage value evaluation. Taking heritage sites along the Yellow River in Shaanxi as a case study, we construct a ten-dimensional value system using Latent Dirichlet Allocation (LDA) and establish the first Cultural Heritage Value Recognition (CHVR) dataset. To enable quantitative evaluation, we introduce the Top-k Heritage Value Weighting (TK-HVW) method, which derives value weights by extracting and normalizing probability distributions across value categories. Extensive experiments demonstrate that ViT-HVE outperforms existing state-of-the-art methods, achieving 0.890 Precision, 0.887 F1-Score, and 0.889 Accuracy. Furthermore, the predicted value distributions show strong alignment with expert evaluations, substantiating the reliability of our method and highlighting its potential for cultural heritage value evaluation and decision-making.