A Survey on Causal Representation Learning Techniques to Extract Causal Features for Causal Machine Learning Model Building
摘要
Research on the marriage between causality and machine learning will assist in addressing the fundamental problems and challenges that exist in Traditional Machine Learning (TML) algorithms such as cause-to-effect relationship between spurious correlation variables, accuracy, generalization, explainability, and robustness. Developing a causal-aware machine learning model or training a machine learning (ML) model on a dataset with features with highly causal-effect relationship has an impact on the identified problems and challenges. In this survey, a review of Causal Representation Learning (CRL) techniques for feature extraction on tabular data was done with the aim of providing a comprehensive and structured analysis of the existing research on causal representation learning techniques for extracting causal features and compared them on their performance, strength, and limitations in building causal machine learning models with minimum generalization error on unseen data. The CRL techniques which were identified and compared on their performance, strength and limitations are Causal Inference Networks (Deep Learning models), Bayesian Networks (BN), Causal Tree Ensembles (CTE), Causal Feature Learning (CFL), Structural Equation Model (SEM), Invariant Causal Prediction and Counterfactual Reasoning Methods (CRM). The study recommends a causal representation learning technique application in real-world observational tabular data to improve the causal machine learning model generalization on unseen data.