Who Knocks on Heaven’s Door: Measuring Augmentation and Outperformance in Human–AI Diagnostic Teams
摘要
Human–AI collaboration in medical diagnostics offers significant potential, yet its effectiveness remains debated. This study investigates two key phenomena in hybrid intelligence systems: human augmentation, where AI improves user performance, and outperformance, where the human–AI team exceeds both individual components. In a multi-site study involving 330 doctors and 16,641 cases across six diagnostic modalities, participants consulted an AI with 81% accuracy during diagnostic tasks. Post-consultation accuracy increased from 75% to 79%, with 57% of participants improving. Outperformance occurred in 18% of cases overall and 27% among those initially less accurate than the AI. The largest gains were observed among lower performers, while 11% of participants—mostly those initially more accurate than the AI—saw declines. These results support the potential of well-integrated AI to enhance diagnostic accuracy and highlight the importance of fostering calibrated trust and effective interaction design to realize the benefits of hybrid intelligence while mitigating risks of over-reliance.