Knowledge Bases, Datasets, and Evaluation
摘要
In this chapter, we delve into the data and evaluation methodologies associated with the entity discovery and linking problem (EDL). A crucial input for this task is the target knowledge base. Some knowledge bases feature rich textual descriptions of entities, complemented by a plethora of attributes and relationships among these entities. In contrast, others might simply provide a brief description for each listed entity. There are knowledge bases that encompass entities articulated in multiple languages, while some are tailored to a specific language or domain. The content within a knowledge base dictates the challenges faced when anchoring entities to varying bases. In addition to discussing mainstream knowledge bases, this chapter will also touch upon training and test data. The quality and volume of annotated data often steer the selection of machine learning models and their subsequent performance. It’s noteworthy that datasets might operate under varied assumptions and present distinct definitions of the problem at hand. Grasping the intricacies of these resources is pivotal, ensuring their optimal utilization and enabling a nuanced comparison of research findings.