Background <p>Evaluating the performance of predictive models for survival is essential before they can be trusted for real-world applications and decision making. While good measures such as the C-index are available for model discrimination, the toolbox for model calibration is much more limited in the time-to-event setting. </p> <p>The method of D-calibration was therefore an important contribution that yields a single numeric value for calibration across the available follow-up time. D-calibration consists of performing a Pearson’s goodness-of-fit test on transformed survival times. Censored survival times are handled using an imputation approach which however tends to yield a conservative test and loss of power.</p> Methods <p>In this paper, we introduce A-calibration based on Akritas’s goodness-of-fit test which is designed specifically for censored time-to-event data. Through theoretical arguments, simulations, and a case study, we compare A- and D-calibration as measures of calibration. In the simulation study, the power of each test to reject a false null-hypothesis was assessed for varying censoring mechanisms (memoryless, uniform and zero censoring), censoring rates, and parameter values of the predictive model considered.</p> Results <p>The simulation study demonstrated that A-calibration had similar or superior power to D-calibration in all considered cases, and that D-calibration, unlike A-calibration, was particularly sensitive to censoring.</p> Conclusions <p>Advantages of A-calibration compared to D-calibration have been demonstrated through theoretical considerations, a simulation study, and a case study, while no disadvantages relative to D-calibration were identified.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A-calibration: assessment of prediction models for survival data under censoring

  • Mikkel Runason Simonsen,
  • Rasmus Plenge Waagepetersen

摘要

Background

Evaluating the performance of predictive models for survival is essential before they can be trusted for real-world applications and decision making. While good measures such as the C-index are available for model discrimination, the toolbox for model calibration is much more limited in the time-to-event setting.

The method of D-calibration was therefore an important contribution that yields a single numeric value for calibration across the available follow-up time. D-calibration consists of performing a Pearson’s goodness-of-fit test on transformed survival times. Censored survival times are handled using an imputation approach which however tends to yield a conservative test and loss of power.

Methods

In this paper, we introduce A-calibration based on Akritas’s goodness-of-fit test which is designed specifically for censored time-to-event data. Through theoretical arguments, simulations, and a case study, we compare A- and D-calibration as measures of calibration. In the simulation study, the power of each test to reject a false null-hypothesis was assessed for varying censoring mechanisms (memoryless, uniform and zero censoring), censoring rates, and parameter values of the predictive model considered.

Results

The simulation study demonstrated that A-calibration had similar or superior power to D-calibration in all considered cases, and that D-calibration, unlike A-calibration, was particularly sensitive to censoring.

Conclusions

Advantages of A-calibration compared to D-calibration have been demonstrated through theoretical considerations, a simulation study, and a case study, while no disadvantages relative to D-calibration were identified.