Semantic scene understanding in computer vision image/video analysis involves assigning semantic labels to each pixel in an image. The main aim of this paper is to partition an image into meaningful scenes and objects and mapping to a specific class. For semantic scene understanding and segmentation, the encoder-decoder approach with UNet is presented in this paper. Encoder network down-sample the image feature and decoder network up-sample the feature map to obtain segmentation mask. The segmentation model is developed that is trained on public Standford dataset images consisting of objects and scenes. The segmented classes used are building, tree, sky, road, grass, river, mountain and foreground objects. The result shows that objects and scenes in the images are partitioned and classified into different classes with an average accuracy of 76.83%. The significance of this research is to recognize, partition and classify the objects and scenes for understanding the segmented contents in video. It can be further used for content-based video retrieval, browsing and summarization. Researchers, Government agencies and Automation industries will benefit from this study.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semantic Segmentation of Natural Scene Images Using Encoder-Decoder Approach

  • K. C. Hari,
  • Manish Pokharel,
  • Sushil Shrestha

摘要

Semantic scene understanding in computer vision image/video analysis involves assigning semantic labels to each pixel in an image. The main aim of this paper is to partition an image into meaningful scenes and objects and mapping to a specific class. For semantic scene understanding and segmentation, the encoder-decoder approach with UNet is presented in this paper. Encoder network down-sample the image feature and decoder network up-sample the feature map to obtain segmentation mask. The segmentation model is developed that is trained on public Standford dataset images consisting of objects and scenes. The segmented classes used are building, tree, sky, road, grass, river, mountain and foreground objects. The result shows that objects and scenes in the images are partitioned and classified into different classes with an average accuracy of 76.83%. The significance of this research is to recognize, partition and classify the objects and scenes for understanding the segmented contents in video. It can be further used for content-based video retrieval, browsing and summarization. Researchers, Government agencies and Automation industries will benefit from this study.