<p>The extraction of road surfaces and centerlines from the satellite images is the most challenging task in the feature extraction field. In recent years, road extraction has been used for diverse applications such as rescue missions, urban planning, automated map updates and autonomous driving. The previous researches have road region misidentification during complex road conditions, inefficient segmentation of road regions and binary semantic segmentation issues. To address these issues, the novel model named as Pretrained Network with Vision Transformer (Conv-ViT) based encoder-decoder model is proposed that is responsible for predicting accurate road extraction. For road extraction, the satellite images are collected from different datasets and preprocessing is involved in this paper to reduce overfitting issues, eliminate artifacts and avoid distortions present in the images. The Conv-ViT model extracts contextual features to produce feature representation. The high and low-level features in the preprocessed images are extracted by using Inception-V3 and ResNet-50. The relationships between the nearby pixels and long-distance pixels are developed by leveraging the ViT transformer model. The encoder captures the generated feature representation and produces the road extraction output from the extracted feature representation. The Mutation Boosted Whale optimization Algorithm (MBWOA) approach is included in this paper that fine-tunes the parameters of the Conv-ViT model. The comparative analyses are performed with the help of performance evaluation measures and attained the highest accuracy rate of 98.7% from the proposed road extraction model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pretrained Network with Vision Transformer for Enhanced Road Extraction from Satellite Imagery

  • K. Madhan Kumar,
  • A. Velayudham

摘要

The extraction of road surfaces and centerlines from the satellite images is the most challenging task in the feature extraction field. In recent years, road extraction has been used for diverse applications such as rescue missions, urban planning, automated map updates and autonomous driving. The previous researches have road region misidentification during complex road conditions, inefficient segmentation of road regions and binary semantic segmentation issues. To address these issues, the novel model named as Pretrained Network with Vision Transformer (Conv-ViT) based encoder-decoder model is proposed that is responsible for predicting accurate road extraction. For road extraction, the satellite images are collected from different datasets and preprocessing is involved in this paper to reduce overfitting issues, eliminate artifacts and avoid distortions present in the images. The Conv-ViT model extracts contextual features to produce feature representation. The high and low-level features in the preprocessed images are extracted by using Inception-V3 and ResNet-50. The relationships between the nearby pixels and long-distance pixels are developed by leveraging the ViT transformer model. The encoder captures the generated feature representation and produces the road extraction output from the extracted feature representation. The Mutation Boosted Whale optimization Algorithm (MBWOA) approach is included in this paper that fine-tunes the parameters of the Conv-ViT model. The comparative analyses are performed with the help of performance evaluation measures and attained the highest accuracy rate of 98.7% from the proposed road extraction model.