Toward a Human-in-the-Loop Approach to Create Training Datasets for RDF Lexicalisation
摘要
Datasets that include alignments between natural language and Knowledge Graphs are fundamental to a wide variety of Natural Language Processing and Generation tasks. Current state-of-the-art aligned datasets, though, are significantly impacted by reduced size and scarcity of covered domains, and their quality is difficult to evaluate. To compensate for these issues, we introduce SEALIon, a tool for extracting RDF triples from natural language textual corpora based on a human-in-the-loop approach. We present our first results of SEALIon’s approach, paving the way for further researches in the field of human-in-the-loop triple extraction.