Causal effect estimation for ordinal outcomes: a performance comparison of G-computation, IPW, and TMLE-based methods
摘要
Ordinal outcomes are frequently used in clinical and observational studies. However, because ordinal variables do not have clearly defined intervals between categories, causal effect estimation for ordinal outcomes is more challenging than that for binary or continuous outcomes. Therefore, accurate causal inference for ordinal outcomes is important, as inadequate covariate adjustment may lead to biased estimates due to confounding. In this study, we defined the average treatment effect (ATE) for ordinal outcomes and compared the performance of various causal inference methods for estimating the ATE.
MethodsA total of nine methods were compared, including the adjusted estimator, parametric- and Super Learner–based G-computation, inverse probability weighting (IPW), and both threshold-based and continuous-score Targeted Maximum Likelihood Estimation (TMLE). Various simulation scenarios were constructed according to whether the outcome and treatment models were correctly specified. The methods were evaluated in terms of bias, empirical standard error (ESE), and root mean squared error (RMSE). In the real-data analysis, the Primary Biliary Cirrhosis (PBC) dataset was used to estimate the average treatment effect of D-penicillamine treatment, and the results were compared across methods using confidence intervals. Confidence intervals were constructed using model-based, bootstrap-based, or influence function–based variance estimation according to the characteristics of each estimator.
ResultsSimulation results showed that, under the specific data-generating mechanisms considered in this study, TMLE-based methods generally exhibited the most stable overall performance. In particular, threshold-based TMLE maintained relatively low bias and RMSE under single-model misspecification scenarios. In contrast, G-computation– and IPW-based methods were sensitive to misspecification of the outcome and treatment models, respectively, and IPW-based methods exhibited substantial variability. When both models were simultaneously misspecified, the performance of most methods deteriorated. The Super Learner library, consisting of a generalized linear model and a random forest, improved the performance of several outcome model–based methods but did not fully mitigate the variability observed in IPW-based methods. For TMLE-based approaches, the Super Learner–based and parametric implementations showed comparable performance.
In the real-data analysis, most methods yielded negative treatment effect estimates, suggesting a potential reduction in disease severity associated with treatment. However, because most 95% confidence intervals included zero, evidence for a statistically significant treatment effect was limited. Methods based on the Super Learner library generally produced narrower confidence intervals, suggesting more stable estimation performance. Nevertheless, differences in treatment effect estimates across methods were small, which may reflect the limited role of confounding in this randomized clinical trial dataset.
ConclusionsUnder the specific data-generating mechanisms considered in this study, threshold-based TMLE demonstrated relatively stable performance for estimating the average treatment effect with ordinal outcomes, particularly under single-model misspecification scenarios. Threshold-based TMLE implemented using the Super Learner library performed comparably to its parametric counterpart. In the real-data analysis, methods based on the Super Learner library tended to produce smaller standard errors and narrower confidence intervals, suggesting the potential for more stable estimation in this setting.