错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A comparative analysis of three independent studies to validate a case difficulty construct for video-based assessment (VBA)

  • Gina L. Adrales,
  • Francesco Ardito,
  • Pradeep Chowbey,
  • Alberto R. Ferreres,
  • Chrys Hensman,
  • Hanno Matthaei,
  • Salvador Morales-Conde,
  • Keith J. Roberts,
  • Harald Schrem,
  • Eric Vibert

摘要

Purpose

A resident’s technical proficiency may vary based on pathology and anatomy. This study was designed to evaluate a simple global case difficulty (GCD) scoring rubric as an important step in defining stages of proficiency required to grant autonomy in the era of entrustable professional activities (EPAs).

Methods

During the implementation of video-based assessment studies to develop OPSAs for laparoscopic cholecystectomy (LC), Roux-en-Y gastric bypass (RYGB), and inguinal hernia repair (IHR), board-certified surgeons were asked to rate GCD as “easy” (1), “average” (2), or “hard” (3). A total of 660 reviews of a convenience sample of de-identified videos were completed: LC: 35 cases by 12 surgeons; RYGB: 30 cases by 4 surgeons; and IHR: 20 cases by 6 surgeons. Raters were instructed not to score based on a general impression of surgeon skill, technique, or instrumentation. Inter-rater reliability was measured by Gwet’s AC2.

Results

Most LC cases were rated as easy (42.2%) while 85.0% of RYGB and 57.1% of IHR cases were rated as average. As measured by Gwet’s AC2, inter-rater reliability was good for LC (0.67) and IHR (0.61) and excellent for RYGB (0.91). LC had the greatest number of hard cases (24.1%) while RYGB had the lowest (4.2%). At least one polar-opposite GCD score (rated as “easy” and as “hard” by different raters) was present in 13 of 35 (37.1%) LC videos, and 7 of 20 (35%) IHR videos. No polar-opposite ratings were present in the RYGB study. Removing three (LC) and two (IHR) raters with outlier scores reduced the percentage of polar-opposite ratings by 62% for LC and 57% for IHR. Analysis of LC study narrative comments regarding “hard” cases highlights several areas of importance to raters including appearance of anatomy (normal vs. aberrant), presence of inflammation (acute vs. chronic), ability to visualize structures, and specimen size (e.g., large gallstones, large gallbladder).

Conclusions

These findings suggest that a simple GCD instrument has the potential to support scalable surgical competency assessments while identifying important areas for improvement.