The question of whether two images depict the same 3D surfaces, and to what degree, often necessitates an extensive and costly search across various scales, involving the matching and geometric confirmation of numerous local features. This cost multiplies significantly when comparing a query image to an entire gallery, as is the case in visual re-localization. Although geometric verification remains crucial, we introduce an understandable cuboid embedding method that simplifies the scale space search to a straightforward lookup process. Our method assesses the asymmetric correlation between a pair of images. Utilizing training samples with known overlaps of visible 3D surfaces, the network acquires a scene-specific criterion of similarity. This allows us to swiftly determine, for instance, which test image represents a zoomed-in version of another and the specific magnification level applied. Consequently, it suffices to detect local features solely at the designated scale. We show the effectiveness of our network tailored for specific scenes by illustrating how this embedding yield comparable results in image matching, all the while maintaining simplicity, speed, and human comprehensibility.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Forecasting Surface Intersection of Images Utilizing Understandable Cuboid Embeddings

  • Zongcheng Zuo,
  • Yuanxiang Li,
  • Yu Zhou,
  • Tongtong Zhang,
  • Zeming Zhou

摘要

The question of whether two images depict the same 3D surfaces, and to what degree, often necessitates an extensive and costly search across various scales, involving the matching and geometric confirmation of numerous local features. This cost multiplies significantly when comparing a query image to an entire gallery, as is the case in visual re-localization. Although geometric verification remains crucial, we introduce an understandable cuboid embedding method that simplifies the scale space search to a straightforward lookup process. Our method assesses the asymmetric correlation between a pair of images. Utilizing training samples with known overlaps of visible 3D surfaces, the network acquires a scene-specific criterion of similarity. This allows us to swiftly determine, for instance, which test image represents a zoomed-in version of another and the specific magnification level applied. Consequently, it suffices to detect local features solely at the designated scale. We show the effectiveness of our network tailored for specific scenes by illustrating how this embedding yield comparable results in image matching, all the while maintaining simplicity, speed, and human comprehensibility.