Reproducibility of methodological radiomics score (METRICS): an intra- and inter-rater reliability study endorsed by EuSoMII
摘要
To investigate the intra- and inter-rater reliability of the total methodological radiomics score (METRICS) and its items through a multi-reader analysis.
Materials and methodsA total of 12 raters with different backgrounds and experience levels were recruited for the study. Based on their level of expertise, raters were randomly assigned to the following groups: two inter-rater reliability groups, and two intra-rater reliability groups, where each group included one group with and one group without a preliminary training session on the use of METRICS. Inter-rater reliability groups assessed all 34 papers, while intra-rater reliability groups completed the assessment of 17 papers twice within 21 days each time, and a “wash out” period of 60 days in between.
ResultsInter-rater reliability was poor to moderate between raters of group 1 (without training; ICC = 0.393; 95% CI = 0.115–0.630; p = 0.002), and between raters of group 2 (with training; ICC = 0.433; 95% CI = 0.127–0.671; p = 0.002). The intra-rater analysis was excellent for raters 9 and 12, good to excellent for raters 8 and 10, moderate to excellent for rater 7, and poor to good for rater 11.
ConclusionThe intra-rater reliability of the METRICS score was relatively good, while the inter-rater reliability was relatively low. This highlights the need for further efforts to achieve a common understanding of METRICS items, as well as resources consisting of explanations, elaborations, and examples to improve reproducibility and enhance their usability and robustness.
Key Points