Predictive averaging and Rubin’s rule-based model pooling to predict survival risk with imputations in the presence of missing patient data: methodology and verification using two case studies and simulations
摘要
Practical estimation of patient risk is often complicated by predictor values being unobserved or missing. In this paper we study the calibration and application of models for patient risk prediction when both the past patient data to which the prediction models are fit contains such missing values, while the future patient records may also be partially observed. Typical applications are often encountered in prognostic observational research or with registry data.
MethodsThis paper presents a detailed study on the use of multiple imputation in prediction, to deal with missing values in both past (calibration data) as well as future patient predictor records. We present the two distinct approaches which may be considered to use imputations to account for the presence of missing values. One method directly averages survival predictions obtained from separate models which are fitted on distinct imputations of the data. The other method first pools the intermediate effect and baseline estimates of these models, before calculation of the predictions from the pooled model. Such pooled estimates could be obtained from Rubin’s rules, for example. Methods are introduced and demonstrated based on two motivating datasets. The application for Cox regression survival modeling is studied in detail. Method performance is verified through cross-validation, a separate validation set and an extensive simulation study.
ResultsAll methods are comparable with respect to bias and Brier score (accuracy) assessment. Major differences are however found between both methods when comparing predicted per-patient risks between repetitions of the fitting procedure with a different set of imputations and when the same number of imputations is used. Predictive averaging is preferable, because the difference between such replicate predictions can be reduced to zero by increasing the number of imputations, which is not the case with the pooled model.
ConclusionsPredictive averaging should be preferred to use of a single pooled model, when using multiple imputations for missing predictor data. Single imputation should not be used in prediction on missing data.