<p>Automated classification of natural landscapes from images is a new and active research area. It is important research domain due to its vast applications in environmental monitoring, urban planning, and digital art. The review of existing studies reveal that conventional machine learning methods lack to balance global context modeling with fine-grained texture recognition across diverse scene types. To address this issue, in this research study, we propose an ensemble framework that integrates an advanced Vision Transformer (ViT) global encoder with an MLP‐based local feature extractor based on an adaptive gating mechanism. The extensive empirical analysis consists of three phases: (a) multi-classification of 12,000 high‐resolution images, (b) application of eXplainable AI (XAI) techniques for interpretation and (c) statistical tests showing confidence level of the findings. The results show that the proposed ensemble framework achieves a peak classification accuracy of 97.29% and an AUC of 0.97, outperforming state of the art ViT variants such as ConvNeXt, PvTv2, and DeiT. Furthermore, latest XAI Techniques (Grad‐CAM, SHAP, LIME) visualize the model’s performance and validate its reliance on semantically meaningful regions. Statistical tests (t-test, ANOVA, chi-square) on image‐derived features and per‐class accuracies confirm the significance of the results at the 98% confidence level, demonstrate that fusing global self-attention with localized MLP representations in dynamic environmental contexts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Global attention and local features using deep perceptron ensemble with vision Transformers for landscape design detection

  • Sang Guodong,
  • Wang Huiyu

摘要

Automated classification of natural landscapes from images is a new and active research area. It is important research domain due to its vast applications in environmental monitoring, urban planning, and digital art. The review of existing studies reveal that conventional machine learning methods lack to balance global context modeling with fine-grained texture recognition across diverse scene types. To address this issue, in this research study, we propose an ensemble framework that integrates an advanced Vision Transformer (ViT) global encoder with an MLP‐based local feature extractor based on an adaptive gating mechanism. The extensive empirical analysis consists of three phases: (a) multi-classification of 12,000 high‐resolution images, (b) application of eXplainable AI (XAI) techniques for interpretation and (c) statistical tests showing confidence level of the findings. The results show that the proposed ensemble framework achieves a peak classification accuracy of 97.29% and an AUC of 0.97, outperforming state of the art ViT variants such as ConvNeXt, PvTv2, and DeiT. Furthermore, latest XAI Techniques (Grad‐CAM, SHAP, LIME) visualize the model’s performance and validate its reliance on semantically meaningful regions. Statistical tests (t-test, ANOVA, chi-square) on image‐derived features and per‐class accuracies confirm the significance of the results at the 98% confidence level, demonstrate that fusing global self-attention with localized MLP representations in dynamic environmental contexts.