MetaCoorNet: an improved generated residual network for grasping pose estimation
摘要
Robotic grasping presents significant challenges due to variations in object properties, environmental complexities, and the demand for real-time operation. This study proposes the MetaCoorNet (MCN), which is a novel deep learning architecture specifically designed to address these challenges in robotic grasping pose estimation. By combining spatial and channel operators, the MetaCoor block is utilized to extract features efficiently. This architecture enhances feature selectivity by embedding location information into channel attention using a positional embedding technique within the coordinate attention mechanism. Consequently, the proposed MCN can focus on pertinent grasp-related regions. Furthermore, convolutional fusion blocks seamlessly integrate spatial and channel features, resulting in enhanced feature resolution and representation capabilities. This innovative design enables the proposed MCN to achieve state-of-the-art performance on the Cornell and Jacquard datasets, attaining accuracies of 98% and 91.2%, respectively. The effectiveness and robustness of MCN are further validated through real-world experiments conducted using a seven-degree-of-freedom Kinova manipulator.