<p>Most previous approaches to thermal-visible video registration generally rely on strict feature point matching and struggle with complex scenes. This paper proposes motion representation-based homography estimation (MRHE), a new video registration approach that leverages end-to-end unsupervised dense optical flow matching. First, we represent video registration clues using optical flow and assess their significance. Then, we construct a homography estimation network based on dense optical flow to estimate the homography matrix. Finally, an exponential smoothing algorithm is introduced to update the homography matrix for stable registration across frames. MRHE offers several advantages: (1) high learning efficiency through an end-to-end unsupervised learning manner, (2) high accuracy and robustness for both stationary and moving platforms, and (3) adaptability to depth variations across frames. We construct a thermal-visible video dataset comprising 500 video pairs across 12 scenes. Experimental results demonstrate that MRHE achieves performance close to the ground truth.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

End-to-end unsupervised optical flow matching for thermal-visible video registration

  • Guohua Feng,
  • Youtian Du,
  • Kang Tian,
  • Yu Lan,
  • Ying Xu,
  • Zhe Hou

摘要

Most previous approaches to thermal-visible video registration generally rely on strict feature point matching and struggle with complex scenes. This paper proposes motion representation-based homography estimation (MRHE), a new video registration approach that leverages end-to-end unsupervised dense optical flow matching. First, we represent video registration clues using optical flow and assess their significance. Then, we construct a homography estimation network based on dense optical flow to estimate the homography matrix. Finally, an exponential smoothing algorithm is introduced to update the homography matrix for stable registration across frames. MRHE offers several advantages: (1) high learning efficiency through an end-to-end unsupervised learning manner, (2) high accuracy and robustness for both stationary and moving platforms, and (3) adaptability to depth variations across frames. We construct a thermal-visible video dataset comprising 500 video pairs across 12 scenes. Experimental results demonstrate that MRHE achieves performance close to the ground truth.