Simulating Causal Data and Evaluation Metrics
摘要
The core challenge in causal inference [1] is the problem of missing counterfactuals. In any real-world dataset, for each individual, we observe only the outcome under the treatment they received. We cannot observe what would have happened had the individual received a different treatment. This problem is sometimes called The Fundamental Problem of Causal Inference. Because of this, it becomes tough to evaluate how well a causal model is performing. How can we check whether our model’s estimates are correct if we do not know the true causal effects? A critical solution to this dilemma is to create simulated datasets, where we, the data creators, control the entire data-generating process. In simulated data, we know both the factual and counterfactual outcomes because we define them. This enables us to measure how close a model’s estimates are to the true causal effects.