WPIML: Web Page Indexing Using Heterogeneous Multi-source Knowledge and Reasoning by Fact Driven Learning
摘要
The epoch of Web 3.0, there arises a necessity for a strategic framework that focuses on a knowledge-centric strategy for indexing web pages. Existing web page indexing frameworks predominantly rely on either keywords or clustering mechanisms. The current demand is for a knowledge-centric model that incorporates learning and reasoning. This paper introduces a statistical model for web page indexing, which not only involves various sources of knowledge but also contributes significantly to web page indexing. The framework's operational process involves the extraction of URLs and descriptions from the dataset. Furthermore, keywords from the URLs are extracted, and an RDF generation paradigm is implemented. This is coupled with ontology generation and knowledge incorporation, utilizing knowledge stores like Wikidata and YAGO. The substantial augmentation of knowledge is accomplished through an ingenious metadata generation strategy. To categorize the metadata, a deep learning Convolutional Neural Network (CNN) classifier is integrated. Additionally, a XGBoost classifier is employed for web page indexing, with the selection of features being driven by Shannon's entropy mechanism. The model facilitates strategic reasoning via normalized Google distance and Renkonen index, employing differential ratios of step deviance measures. The optimization prowess of the artificial algae algorithm is harnessed for effective solutions. This optimization aids in achieving the finest instances for web page indexing, thereby narrowing the cognitive gap between the existing worldwide web and the incorporated knowledge of Web 3.0. The proposed framework's performance is remarkable, boasting an average precision of 96.45%, an average recall of 97.19%, and an impressive F-measure of 387.27%. The framework also achieves the lowest value of False Discovery Rate (FDR) at a mere 0.04%.