Estimating textual treatment effect via causal disentangled representation learning
摘要
Estimating causal effects from observational data reveals the potential outcomes of different treatments. However, the methods primarily focusing on numerical or categorical covariates leave causal inference with textual observational data as an unresolved issue. Specifically, the high-dimensional and unstructured nature of text complicates the learning of representation vectors of causal structure from textual covariates. This complexity is principally due to the interaction of different factors within the textual covariates, making separating these factors crucial for the accurate estimation of textual causal effects. To address this challenge, we propose a causal disentangled representation learning method based on variational inference. The method derives latent factors from observed textual covariates and decomposes them into instrumental, confounding, and adjustment factors. Additionally, a learning criterion that minimizes mutual information is employed to ensure the independence of disentangled factors, and targeted regularization based on nonparametric estimation is applied to reduce residual bias. Experimental results show that the proposed method performs well on textual causal effect datasets and has higher performance and competitiveness compared to strong baseline methods.