SemanticDifference: Change Detection with Multi-scale Vision-Language Representation Difference
摘要
Existing change detection methods in 3D scenes are easily affected by lightning and shadows due to image- or CNN feature-level difference lacks of high-level semantic information. In this paper, we propose a novel change detection method, SemanticDifference, which distinguishes the changes from the differences of multi-scale vision-language representation. Our model adopts 4D Gaussian to represent the 3D scenes of the pre-change and post-change scenes. Then we render a post-change image from the same viewpoint of a pre-change image. To capture the changes, we generate multi-scale segments and compare their vision-language embeddings. We also propose a simple strategy to determine whether a changing object appears or disappears. By this, we can generate a change caption of the image pair. Through extensive experiments on various instance-level change detection scenes, our method demonstrates significant improvements in detection accuracy over state-of-the-art methods like C-NeRF and Gaussian Difference, particularly in scenarios with large lighting variations.