This chapter addresses literature-based discovery focused on outlier documents, aiming to improve b-term search efficiency by reducing the search space of potential b-terms to those appearing in outlier documents. Section 6.1 provides a motivation for using outliers in LBD. The next two sections describe two methods for b-term detection in outlier documents. Section 6.2 presents the classification approach for detecting outlier documents, while Section 6.3 outlines a clustering approach to outlier document detection. In Section 6.4, outlier detection is compared to a heuristic approach to b-term detection, as implemented in the CrossBee LBD system. Then, the combined methodology is presented in Section 6.5 as a two-step process combining outlier detection and cross-domain term exploration using the CrossBee tool, illustrated by workflows in the TextFlows Web-based platform. Section 6.6 illustrates the application of the method to the domains of Alzheimer’s disease and gut microbiota. Section 6.7 evaluates and compares the obtained results with other research results and provides a summary and directions for further work. The chapter concludes with Section 6.8, which includes a Python tutorial allowing for simplified LBD experiment replicability and code reuse.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Outlier-based Closed Discovery

  • Nada Lavrač,
  • Bojan Cestnik,
  • Andrej Kastrin

摘要

This chapter addresses literature-based discovery focused on outlier documents, aiming to improve b-term search efficiency by reducing the search space of potential b-terms to those appearing in outlier documents. Section 6.1 provides a motivation for using outliers in LBD. The next two sections describe two methods for b-term detection in outlier documents. Section 6.2 presents the classification approach for detecting outlier documents, while Section 6.3 outlines a clustering approach to outlier document detection. In Section 6.4, outlier detection is compared to a heuristic approach to b-term detection, as implemented in the CrossBee LBD system. Then, the combined methodology is presented in Section 6.5 as a two-step process combining outlier detection and cross-domain term exploration using the CrossBee tool, illustrated by workflows in the TextFlows Web-based platform. Section 6.6 illustrates the application of the method to the domains of Alzheimer’s disease and gut microbiota. Section 6.7 evaluates and compares the obtained results with other research results and provides a summary and directions for further work. The chapter concludes with Section 6.8, which includes a Python tutorial allowing for simplified LBD experiment replicability and code reuse.