I-rod: an ensemble of CNNs for object detection in unconstrained road scenarios
摘要
Solving the problem of object detection in complex and unstructured environments is crucial for enhancing the safety and efficiency of autonomous systems. This paper introduces a semantic segmentation model capable of accurate object detection in complex backgrounds by integrating multiple Convolutional Neural Networks (CNNs). The system incorporates two distinct segmentation models: an Encoder-Decoder architecture for acquiring abstract feature representations and a dilated convolutional branch to tackle variations in object sizes. The model employs a dynamic fusion mechanism based on confidence scores from each branch, allowing it to adapt to varying and dynamic situations. The model is evaluated on the Indian Driving Dataset (IDD), featuring unstructured road environments, and the Cityscape dataset. Comparative pixel-wise analysis shows the proposed model outperforming four other state-of-the-art segmentation models by