Multi-head self-attention convolution and adaptive feature fusion network for pixel-level multi-object planar grasping detection
摘要
Planar grasping detection is to find a set of robotic grasping configurations from top to bottom. Most recent methods use sparse rectangular representation and convolution to generate grasping postures, which may miss the optimal grasp posture and is more prone to grasp local features; these problems may result in relatively poor performance in multi-object grasping task, involving special difficulties of mutual interference, global grasp configuration, and unknown objects. Aiming to address these problems, we propose the novel multi-head self-attention convolution and adaptive feature fusion network (MSACAFF-Net) for pixel-level multi-object planar grasping detection. Firstly, the window multi-head self-attention & big kernel conv network (WMSABC-Net) is proposed to effectively extract and fuse local and global features in a local-to-global attention mechanism; the multi-head self-attention convolution model is to extract local features in windows, and the big kernel convolution model is to extract global features; this local–global information flow extraction process can enhance the ability to extract both local and global information. Then, the adaptive feature fusion network (AFF-Net) is proposed to enhance the correlation between features and grasping parameter space; the dilated convolution is used to generate features with different scales and receptive fields; the features are adaptively fused and enhanced according to the importance to grasping detection, suppressing irrelevant features and enhancing grasp-related features. Experiments show that the MSACAFF-Net can effectively do grasping detection for known and unknown multi-object in the PLGP dataset and Cornell dataset, and achieve State-of-The-Art performance.