Monocular Stereo Learning Based on Hybrid Attention Transformer
摘要
Monocular stereo matching estimation network is an important research direction in the field of artificial intelligence. In this paper, a monocular stereo learning network based on Hybrid Attention Transformer for joint estimation of optical flow and depth for monocular sequences is proposed. Firstly, spatial window and channel group attention are used to enhance the ability of the network to extract local and global features, optimizing the learning framework. Secondly, cross attention is used to realize the cross-window connection, which makes the details of image estimation more rich. Finally, we discuss the effect of the two different connection modes of attention in series and parallel on the network estimation tasks. Our experiment results on KITTI datasets show that, compared with the previous methods, our multi-task learning framework has more significant improvements in optical flow and depth.