Development and Evaluation of a Diagnostic Exam for Undergraduate Biomedical Engineering Students Using GPT Language Model-Based Virtual Agents
摘要
The creation of a diagnostic exam for biomedical engineering undergraduate students using virtual agents based on GPT language models is analyzed in this work. Thirty-nine eighth-semester students answered a 20-question exam generated by ChatGPT-3 covering the topics of acquisition, amplification, processing, and visualization of biomedical signals encompassing different levels of thinking according to the taxonomy of Bloom, including the application, analysis, and evaluation levels. Three academic experts assessed the quality of the questions based on clarity, relevance, level of thinking, and difficulty. Also, difficulty and discrimination indexes and Rasch analysis were calculated. Students obtained an average grade of 5.91, with a standard deviation of 1.39 points. Subject reliability was 0.599, and the p-value for the fit of the model of Rasch was 0.017. High correlations between some questions were observed. Based on their difficulty, a few questions could be considered irrelevant. The Wright competency map showed a good distribution on the ability scale with some redundancies and gaps. In conclusion, virtual agents have great potential to create diagnosis exams in biomedical engineering. However, it is necessary to consider their limitations and conduct a rigorous evaluation of the quality and reliability of the questions generated.