Unbiased Image Caption Generation Based on Dynamic Counterfactual Inference
摘要
Recent research has focused on exploring causality in image description generation. Some studies have proposed using backdoor criteria to eliminate visual confounding factors and language confounding factors in images and annotations of datasets, by computing co-occurrence probabilities between objects and between words. These methods can effectively deconfound the visual and language confounders simultaneously. However, the impact of language priors on generating descriptions still needs to be explored during sequence generation. Language priors can be classified into two types: good and bad. Good language priors make word generation more fluent, while bad language priors can lead to spurious associations between generated objects. Therefore, in this paper, we first investigate how image caption generation models introduce bad language priors during the word generation process. Inspired by the causal effect, we propose a framework based on dynamic counterfactual inference to eliminate the generation of false objects. This framework utilizes noisy features and input words to capture the direct language effect during word generation, and reduces the influence of bad language bias by dynamically subtracting the direct language effect. Experimental results demonstrate that our proposed counterfactual inference framework achieves competitive performance on the MSCOCO dataset for image caption generation tasks, and effectively reduces the incidence of false objects in the description.