Identifying Relevant Data in RDF Sources
摘要
The increasing number of RDF data sources published on the web represents an unprecedented amount of information. However, querying these sources to extract the relevant information for a specific need represented by a target schema is a complex task as the alignment between the target and the source schemas might not be provided or incomplete. This paper presents an approach which aims at automatically populating the classes of a target schema. Our approach relies on a semi-supervised learning algorithm that iteratively identifies instance patterns in the data source that represent candidate instances for the target schema. We present some preliminary experiments showing the effectiveness of our approach.