Using bandit algorithms to design response-adaptive trials can optimize participant outcomes, but poses major challenges for statistical inference. Recent attempts to address these challenges typically impose restrictions on the exploitative nature of the bandit algorithm and require large sample sizes to ensure asymptotic guarantees. However, large experiments generally follow a successful pilot study, which is tightly constrained in its size or duration. In this work, we tackle the problem of hypothesis testing in finite samples. We illustrate an innovative hypothesis testing procedure, uniquely based on the allocation probabilities of the bandit algorithm, and theoretically characterise it when applied to Thompson sampling.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Finite-Sample Inference in Response-Adaptive Designs: An Application to Thompson Sampling

  • Nina Deliu,
  • Joseph J. Williams,
  • Sofia S. Villar

摘要

Using bandit algorithms to design response-adaptive trials can optimize participant outcomes, but poses major challenges for statistical inference. Recent attempts to address these challenges typically impose restrictions on the exploitative nature of the bandit algorithm and require large sample sizes to ensure asymptotic guarantees. However, large experiments generally follow a successful pilot study, which is tightly constrained in its size or duration. In this work, we tackle the problem of hypothesis testing in finite samples. We illustrate an innovative hypothesis testing procedure, uniquely based on the allocation probabilities of the bandit algorithm, and theoretically characterise it when applied to Thompson sampling.