Natural Language Processing of Structured Text Elicitations Around the Replicability of Scientific Claims
摘要
How do experts talk about the replicability of claims in the social and behavioral sciences (SBS)? In support of a DARPA program looking at replicability in SBS, RAND was asked to explore whether the way experts discussed claims could aid in the design of explainable outputs from Artificial Intelligence (AI) systems that estimated the likelihood that research claims would be replicable. To answer this question, RAND analyzed the online discourse of experts participating in a Delphi process conducted by the University of Melbourne that evaluated thousands of claims published in disciplinary and multidisciplinary SBS research. We analyzed the expert discourse by developing several Natural Language Processing (NLP) models that associated different features of speech and text (words, phrases, and stance categories) with expert assignments of replicability scores to claims. We found that the way in which experts talked about a given claim, i.e., features of linguistic stance, was a significant predictor of the replicability scores experts gave, separately from the substance of their evaluation, i.e., the emphasis on specific experimental features, such as sample size or assignment of treatments to study participants. Moreover, linguistic style’s predictive power regarding replicability scores decreased as the uncertainty of evaluators increased. Overall, the signals we found offer an initial template regarding how scientific claims are discussed in expert discourse and the ways in which experts explain their evaluations. While the development of AI systems that estimate the likelihood that scientific claims may be replicable remains technically ambitious, this analysis offers preliminary results that will help make their outputs understandable to users.