Out-of-distribution detection in text using statistical techniques
摘要
Out-of-distribution data points diverge from the general profile of the data, typically defined by the specific task for which the machine learning model is being constructed. Machine learning models are more reliable when out-of-distribution detection is part of the pipeline. Out-of-domain detection models are employed not just to sieve input into a model, but also to scrutinise output from a generative model, a process known as selective generation [