错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SACFormer: Unify Depth Estimation and Completion with Prompt

  • Shiyu Tang,
  • Di Wu,
  • Yifan Wang,
  • Lijun Wang

摘要

Monocular depth estimation and depth completion are closely correlated, yet they have long been approached as two distinct and separate tasks. In this paper, we propose a new Transformer architecture dubbed SACFormer to unify these two tasks with Spatially Aligned Cross-modality (SAC) attention modules. Unlike existing unified methods, SACFormer is able to take advantage of the correlation between two tasks using a single model and one set of network parameters without task- or modality-specific modules. To better identify their unique characteristics, we further introduce a window-based deep prompt learning scheme that enables SACFormer to seamlessly switch between the two tasks. By integrating the above two contributions, we are able to enforce the synergy between depth estimation and completion, while respecting their differences. As a result, our method using one model trained on multi-domain data can simultaneously handle the two related tasks, thereby significantly reducing memory footprint during deployment. The proposed method is extensively evaluated on popular benchmarks and performs favorably in both indoor and outdoor scenes.