Outlier-based Closed Discovery
摘要
This chapter addresses literature-based discovery focused on outlier documents, aiming to improve b-term search efficiency by reducing the search space of potential b-terms to those appearing in outlier documents. Section 6.1 provides a motivation for using outliers in LBD. The next two sections describe two methods for b-term detection in outlier documents. Section 6.2 presents the classification approach for detecting outlier documents, while Section 6.3 outlines a clustering approach to outlier document detection. In Section 6.4, outlier detection is compared to a heuristic approach to b-term detection, as implemented in the CrossBee LBD system. Then, the combined methodology is presented in Section 6.5 as a two-step process combining outlier detection and cross-domain term exploration using the CrossBee tool, illustrated by workflows in the TextFlows Web-based platform. Section 6.6 illustrates the application of the method to the domains of Alzheimer’s disease and gut microbiota. Section 6.7 evaluates and compares the obtained results with other research results and provides a summary and directions for further work. The chapter concludes with Section 6.8, which includes a Python tutorial allowing for simplified LBD experiment replicability and code reuse.