Measuring Data Sufficiency
摘要
The problem of determining data sufficiency is a repeating issue. An often quoted response to the problem is: “There is no data, like more data” (Pieraccini (There Is No Data like More Data, in ‘The Voice in the Machine: Building Computers That Understand Speech’, The MIT Press, 2012)). Indeed, increasing the sample size is definitely a guaranteed way to increase the probability of coming up with a good model. However, data acquisition can be costly, the task can be time sensitive, or there is simply not more data available. For example, during the COVID19 pandemic, there was a discussion between one of the vaccine providers and the US government of whether the sample size of tested subjects was large enough to warrant approval. (“FDA advisers debate standards on a coronavirus vaccine for young children,” Washington Post, June 10, 2021.) Here the task was time sensitive, data acquisition was costly, and furthermore, the stakes were high. This chapter discusses a method to estimate data sufficiency based on the information metrics earlier in this book. The method clearly shows when there is enough data and also clearly shows when there are not enough samples. It can be ambiguous when there may or may not be enough data. Most importantly, it works without building a model.