Artificial intelligence for automatic FLS credentialing: highlighting and addressing current limitations
摘要
Artificial Intelligence (AI) can automate technical skills assessment. Applied to the Fundamentals of Laparoscopic Surgery (FLS®), AI-driven credentialing could enhance evaluation consistency while reducing costs associated with human proctoring. However, existing studies fall short of demonstrating the reliability required for high-stakes assessments, often due to limited validation and suboptimal modeling strategies.
MethodsA novel AI-based approach is proposed to assess the FLS peg transfer task from video recordings. The system analyzes each video frame to track the state of objects on the board (peg state) and classify user behavior during the task (surgical actions). Multiple AI models are used to generate these predictions, which are then refined using two post-processing techniques designed to enhance temporal consistency and task-specific accuracy. To ensure rigorous validation, we assess individual model performance, as commonly done by existing works, and introduce two new metrics—transfer precision and transfer recall—to measure the system’s ability to replicate proctor-level evaluation.
ResultsThe validation dataset includes 21 full-length videos of peg transfer tasks performed by 11 expert and 10 non-expert users. Individual AI models show good performance, with accuracy > 99% for peg state prediction and f1-score > 78% for action recognition, comparable with existing studies. However, when not using post-processing algorithms, transfer precision and recall only reach 22.86% and 55.56%, respectively, despite the good individual AI model performance. With the proposed post-processing, these metrics improve significantly to 80.44% and 96.51%.
ConclusionThis study underscores the importance of task-specific validation and modeling strategies for developing robust AI systems suitable for automatic credentialing. While focused on the FLS peg transfer task, the proposed framework provides generalizable guidelines for building reliable AI-based assessment tools for both simulated and real surgical environments.
Graphical abstract