Spatiotemporal-Modal Collaborative Calibration for Dynamic Depth Video Enhancement
摘要
Modern depth sensors face limitations including low resolution, narrow FOV, and dynamic interference, leading to missing data and distortions. Existing methods mainly process single-frame depth features, underutilizing temporal cues like motion trajectories and occlusions in sequential frames. To address these limitations, we propose a three-stage Guidance-Calibration-Optimization architecture. Spatio-temporal guidance network employs memory units to process sequential frames and generate initial guided depth maps; depth-dominant network performs multi-modal calibration on the guided depth map, RGB semantics, and raw depth map to suppress modal interference and reinforce reliable features; geometric optimization smooths noise in homogeneous regions while preserving boundaries, improving visual consistency. This design reduces optical flow dependency and facilitates depth-focused temporal modeling. Experiments indicate enhanced depth video quality with improved completeness and accuracy, supporting dynamic scene understanding.