Ears for examinations: a multicentre quasi-experimental evaluation of AI-generated revision podcasts on learning outcomes and retention among medical students
摘要
Medical curricula overload necessitates efficient, self-directed learning strategies that can utilize ‘dead time’ for study. AI-generated audio offers a scalable resource for this purpose, yet evidence regarding its efficacy in low-resource settings is limited. This study evaluated the effectiveness and learner perceptions of AI-generated podcasts on learning outcomes and retention among Indian medical students.
MethodsA multicentre, quasi-experimental, mixed-methods explanatory sequential study was conducted at two geographically distinct medical colleges in India. Third-professional undergraduate MBBS students received six AI-generated, faculty-validated revision podcasts covering high-yield Community Medicine topics. Knowledge was assessed via a validated 60-item MCQ test at baseline (pre-test), immediately post-intervention, and at 30 days (delayed post-test). The intervention and assessments targeted both lower-order (remember, understand) and higher-order (apply, analyse) cognitive domains per Bloom’s Taxonomy. The primary analysis employed a Linear Mixed-Effects Model (LMM) to estimate learning trajectories. Qualitative data from Focus Group Discussions (FGDs) were analyzed using Framework Analysis to explore mechanisms of impact and integrate findings with quantitative outcomes.
ResultsThe analytic cohort comprised 239 participants. Both centers demonstrated statistically significant immediate learning gains (Centre 1: Cohen’s dz = 0.42; Centre 2: Cohen’s dz = 0.55) across all levels of cognitive domain, more so in higher domain. Retention trajectories diverged: Centre 1 showed stable retention at 30 days, while Centre 2 experienced significant knowledge decay. In multivariable analysis, higher episode completion (> 50%) was independently associated with greater learning gain (adjusted β = 5.15, p = 0.005) after controlling for baseline scores and centre. Qualitative integration revealed that while learners valued the tool for portability, high-intensity ‘binge-listening’ at Centre 2 aligned with the observed decay, whereas distributed usage at Centre 1 supported stability. Students identified lack of pauses and AI voice monotony as key design barriers.
ConclusionsThe robust post-test rise validates AI-generated podcasts as an effective tool for rapid revision and pre-exam preparation. While the intervention successfully augmented study time, long-term retention favored distributed over massed practice. Future AI tools should incorporate engineered pauses and improved prosody to optimize cognitive load and sustained engagement.