Semantic Segmentation and Scene Understanding Using RGBD Data
摘要
Semantic segmentation is a crucial task in computer vision, aimed at classifying each pixel in an image into predefined categories. Traditional RGB-based semantic segmentation struggles with accuracy in complex, cluttered environments where objects overlap or vary in distance. This work addresses the challenges using a Microsoft Kinect sensor by performing segmentation on multiple objects in RGBD images, subsequently analyzing their spatial and depth information to determine the interrelations between different objects such as cans, boxes, and cups. This approach aims to enhance overall scene comprehension and object interaction analysis.