Psychometric performance of EQ-5D-5L and SF-6Dv2 in cancer patients
摘要
To evaluate the psychometric properties of EQ-5D-5L and SF-6Dv2 in a group of patients with breast or colorectal cancer. EQ-5D-5L, SF-6Dv2, and QLQ-C30 were completed at baseline and follow-up by patients with breast or colorectal cancer in Quebec, Canada. Ceiling effect was assessed by calculating the percentage of respondents reporting the best health state. Agreement between EQ-5D-5L and SF-6Dv2 was evaluated using intraclass correlation coefficients (ICC) and the Bland–Altman plot. Convergent validity for both instruments was assessed using the Spearman rank correlation coefficient (r) with QLQ-C30 serving as a calibration standard. Known-group validity was evaluated by comparing the scores of patients with different health conditions, while sensitivity was further assessed within these known groups using relative efficiency (RE). Finally, test–retest reliability and responsiveness were tested using ICC and the area under the receiver operating characteristic curve (AUC), respectively. 204 patients were enrolled at baseline and 103 were followed up. No ceiling effect was found for SF-6Dv2 compared to 16.67% for EQ-5D-5L. Agreement between EQ-5D-5L and SF-6Dv2 utility values was moderate (ICC = 0.53). The Bland–Altman plot indicated that the differences between the two instruments tended to increase as the average score decreased, reflecting more severe health states. Correlation between the two instruments and with QLQ-C30 score ranged from low to moderate. Each dimension of EQ-5D-5L and SF-6Dv2 had greater correlations with similar dimensions of QLQ-C30 (r > 0.50), except for two dimensions. SF-6Dv2 was able to identify only a minority of known groups, and the RE for SF-6Dv2 in variables was greater than 1, suggesting higher sensitivity. Both instruments had poor test–retest reliability (ICC < 0.40). AUC for EQ-5D-5L utility score was 0.792, while for SF-6Dv2 it was 0.673. EQ-5D-5L and SF-6Dv2 were found to have good convergent validity measures, but each has its own strengths and limitations. These findings suggest that the choice between two instruments should depend on the specific research and clinical context. Further studies are needed to clarify their relative performance in terms of known-group validity.