Background. In recent years, cyber security user studies have been scrutinized for their reporting completeness, statistical reporting fidelity, statistical reliability and biases. It remains an open question what strength of evidence positive reports of such studies actually yield. We focus on the extent to which positive reports indicate relations true in reality, that is, a probabilistic assessment. Aim. This study aims at quantifying overall strength of evidence in cyber security user studies. Method. Based on 431 coded statistical inferences in 146 cyber security user studies from a published SLR covering the years 2006–2016, we first compute a simulation of the a posteriori false positive risk based on parametrized prior probability, biases and effect size thresholds. Second, we establish the observed likelihood ratios for positive reports. Third, we compute the reverse Bayesian argument on the observed positive reports by computing the prior required for a fixed a posteriori false positive rate. Results. We obtain a comprehensive analysis of the strength of evidence of the field. The simulations show that even in face of well-controlled conditions and high prior likelihoods, only few studies achieve good a posteriori probabilities. Conclusions. This work constitutes a “What if?” analysis, which permits the reader to evaluate the consequences of their assumptions on the state of the field. One may stop short at the bleak conclusion that the strength of evidence of the field leaves something to be desired and that most positive reports are likely false. At the same time, the “What if?” analysis offers a way forward to sensitize researchers to the effects of investigating many relations and incurring biases. It, thereby, allows them to plan better ahead for future studies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Why Most Results of Socio-Technical Security User Studies are False

  • Thomas Groß

摘要

Background. In recent years, cyber security user studies have been scrutinized for their reporting completeness, statistical reporting fidelity, statistical reliability and biases. It remains an open question what strength of evidence positive reports of such studies actually yield. We focus on the extent to which positive reports indicate relations true in reality, that is, a probabilistic assessment. Aim. This study aims at quantifying overall strength of evidence in cyber security user studies. Method. Based on 431 coded statistical inferences in 146 cyber security user studies from a published SLR covering the years 2006–2016, we first compute a simulation of the a posteriori false positive risk based on parametrized prior probability, biases and effect size thresholds. Second, we establish the observed likelihood ratios for positive reports. Third, we compute the reverse Bayesian argument on the observed positive reports by computing the prior required for a fixed a posteriori false positive rate. Results. We obtain a comprehensive analysis of the strength of evidence of the field. The simulations show that even in face of well-controlled conditions and high prior likelihoods, only few studies achieve good a posteriori probabilities. Conclusions. This work constitutes a “What if?” analysis, which permits the reader to evaluate the consequences of their assumptions on the state of the field. One may stop short at the bleak conclusion that the strength of evidence of the field leaves something to be desired and that most positive reports are likely false. At the same time, the “What if?” analysis offers a way forward to sensitize researchers to the effects of investigating many relations and incurring biases. It, thereby, allows them to plan better ahead for future studies.