Using Kane’s validity framework to examine the implications of feedback in simulation-based assessments
摘要
Simulation modalities have been increasingly used within programmatic assessment systems, yet educators typically have not collected and appraised validity evidence to justify such uses. Kane’s validity framework offers a contemporary approach to conducting validation studies of assessment practices. Under the framework, educators collect and appraise validity evidence according to four inferences: the scoring of performance, the generalization of scores to other assessment contexts, the extrapolation of assessment performance to real-world contexts, and the implications or consequences of assessment decisions for learners, educators, programs, patients, and society. We developed a simulation-based echocardiography competence assessment tool (ECAT) and collected validity evidence to evaluate its use as an assessment for learning. We applied Kane’s validity framework to evaluate the utility of the ECAT, with a focus on the implications of the assessment for promoting trainees’ learning.
MethodsWe implemented the ECAT in 2017, collecting simulation-based performance data and subsequent interview data. Fourteen cardiology trainees were assessed using the ECAT by four raters, and their performance was video-recorded. After trainees reviewed their performance videos and feedback, we conducted individual interviews with them and the raters who provided feedback. Directed content analysis generated implications and scoring evidence, and quantitative analyses generated scoring and extrapolation evidence. All evidence was critically appraised to form a validity argument about using ECAT as an assessment for learning.
ResultsParticipants reported that ECAT scores accurately represented trainees'performance, and that the feedback helped identify learning opportunities. Inter-rater reliability was high at ICC = 0.913 (95% CI 0.81–0.97). Participants’ ECAT scores correlated with their end-of-rotation cardiology exam scores (r = 0.66, p = 0.02) and had positive associations with raters’ judgments of the diagnostic quality of their scans, and with their reported numbers of echocardiograms seen, performed, and interpreted.
ConclusionsOur integrated analysis produced a data-informed validity argument supporting the use of the ECAT as a simulation-based assessment for learning. The findings also highlighted multiple areas for further research to optimize the ECAT. Our illustrative example of Kane’s validity framework aims to support simulation educators as they are increasingly called on to justify the use of simulation-based assessments in programmatic and competency-based assessment systems.