Analysis of UNet-Based Semantic Segmentation Models
摘要
Semantic segmentation is one of the prime areas of research for computer vision applications. Many models have been developed by the researchers for various real-world applications. However, these models have rarely been analyzed over a similar set of datasets. Therefore, this paper presents a comprehensive assessment of semantic segmentation models on complex images, exclusively focusing on UNet-based architectures. The study evaluates the analyzed models using five publicly available datasets, namely Oxford-IIIT Pets, CT liver, Martial Arts, Dancing and Sports, Kvasir-Sessile, and Indian Driving Dataset-Lite ensuring a robust and transparent validation process. The findings highlight the significant impact of employing the EfficientNetB7 encoder alongside the UNet architecture, leading to significant improvements in Intersection over Union scores. This enhancement signifies a substantial advancement in scene segmentation capabilities, demonstrating the potential of this fusion approach in advancing computer vision applications like autonomous driving systems and medical imaging.