Deep Learning Method in Object Detection: R-CNN to Mask R-CNN
摘要
Convolutional neural networks (CNNs) are crucial for object detection. It is surpassing traditional, handcrafted feature-based methods. The use of handmade attributes is outpaced by ML-based object detection that learns both high and low level features. Recent research encompasses topics such as classical structures, training methodologies, loss functions, backbone networks, datasets, metrics, applications, and future directions in computer vision. Accurate individual identification in images is crucial for applications like interactive computing and surveillance. Despite progress, precise localization remains challenging especially for complex backgrounds. G-Mask (Guided Attention Mask) enhances person detection in Mask R-CNN, refining accuracy, particularly around object perimeters. Instance segmentation aims for a comprehensive understanding of a person’s spatial distribution. BMask R-CNN, an improved version of Mask R-CNN, employs an extra mechanism to enhance object boundary distinction, improving effectiveness, especially in scenarios demanding precision. Traditional Mask R-CNN may face challenges in precise classification near object boundaries, leading to potential inaccuracies. BMask R-CNN, incorporating instance boundary information, significantly improves mask prediction accuracy, especially crucial for fine boundary detection in challenging scenarios. Extensive testing on various datasets demonstrates its superiority over conventional methods, advancing computer interpretation of visual content, particularly in person and object detection within the Mask R-CNN framework. This research brings us closer to computers adeptly interpreting visual content in nuanced scenarios.