Background <p>When analyzing randomized controlled trials (RCTs) data, covariate adjustment is often employed to increase the precision of estimated treatment effects. Missing data in covariates, if not handled properly, can result in biased and inefficient estimates. However, the existing literature on handling missing covariate data is limited, and recommendations vary regarding a valid and efficient approach.</p> Methods <p>To help reconcile the seemingly inconsistent recommendations, we address two questions through methodological descriptions and simulated demonstrations. First, how should a multiple imputation (MI) model be specified for RCTs to best preserve the benefit of the randomization design? We consider three different approaches: MI with only baseline variables, “MI overall”, and “MI by arm”. Second, when and why will simple general strategies, such as grand mean imputation and the missing indicator method, perform as well as or better than MI in estimating treatment effects, and when and why do they fail?</p> Results <p>“MI by arm” has the potential to produce unbiased estimates for both the average and subgroup treatment effect (primary and secondary analyses) under the missing at random assumption. Strategies that capitalize on the randomization design, including MI with baseline variables, grand mean imputation, and the missing indicator method, may generate unbiased estimates for the average treatment effect (primary analysis) regardless of the missing data mechanism.</p> Conclusion <p>This article clarifies the assumptions and mechanisms by which different missing data strategies accommodate missingness in covariates and reconcile recommendations that sometimes appear contradictory in the literature. Under MAR, “MI by arm” produces unbiased estimates for both the average treatment effect and subgroup treatment effects. Leveraging the randomization design, “baseline-only MI”, grand mean imputation, and the missing indicator method produce unbiased estimates for the average treatment effect, but biased subgroup treatment effects, regardless of the missing data mechanism.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

How to manage missing covariates in randomized controlled trials: a comparison of strategies

  • Shiyu Zhang,
  • Yajuan Si,
  • John J. Dziak

摘要

Background

When analyzing randomized controlled trials (RCTs) data, covariate adjustment is often employed to increase the precision of estimated treatment effects. Missing data in covariates, if not handled properly, can result in biased and inefficient estimates. However, the existing literature on handling missing covariate data is limited, and recommendations vary regarding a valid and efficient approach.

Methods

To help reconcile the seemingly inconsistent recommendations, we address two questions through methodological descriptions and simulated demonstrations. First, how should a multiple imputation (MI) model be specified for RCTs to best preserve the benefit of the randomization design? We consider three different approaches: MI with only baseline variables, “MI overall”, and “MI by arm”. Second, when and why will simple general strategies, such as grand mean imputation and the missing indicator method, perform as well as or better than MI in estimating treatment effects, and when and why do they fail?

Results

“MI by arm” has the potential to produce unbiased estimates for both the average and subgroup treatment effect (primary and secondary analyses) under the missing at random assumption. Strategies that capitalize on the randomization design, including MI with baseline variables, grand mean imputation, and the missing indicator method, may generate unbiased estimates for the average treatment effect (primary analysis) regardless of the missing data mechanism.

Conclusion

This article clarifies the assumptions and mechanisms by which different missing data strategies accommodate missingness in covariates and reconcile recommendations that sometimes appear contradictory in the literature. Under MAR, “MI by arm” produces unbiased estimates for both the average treatment effect and subgroup treatment effects. Leveraging the randomization design, “baseline-only MI”, grand mean imputation, and the missing indicator method produce unbiased estimates for the average treatment effect, but biased subgroup treatment effects, regardless of the missing data mechanism.