Evaluation
摘要
Suppose you are in charge of administering an enterprise database and are tasked with choosing an NLIDB system to provide a non-technical interface to the users of your database. You check the literature and find hundreds of research papers on the subject. You also check a few leaderboards and find tens of top-performing models. How do you decide on a system? What should you look for in choosing one from many choices? These are some of the questions that are answered in this chapter. This chapter covers the evaluation methodologies for NLIDBs with a focus on text-to-SQL. Datasets and benchmarks are presented and their statistics, in terms of supported query types, are compared. Two major evaluation methods, namely reference-based and human- centric evaluations, are discussed in more detail. Finally, other evaluation metrics are reviewed and some of the top-performing models and their characteristics are presented.