Background Technologies
摘要
This chapter presents a selection of the basic technologies used as elementary components of LBD methods. Section 3.1 provides an outline of information extraction methods. Section 3.2 presents a standard natural language processing (NLP) pipeline used in text analysis. Section 3.3 presents the Bag-of-Words representation of text documents and similarity measures used in standard text mining tasks, such as clustering. Section 3.4 introduces the area of network analysis, Section 3.5 presents selected embedding-based analysis approaches, and Section 3.6 briefly introduces the Large Language Model (LLM) technology. The evaluation of performance of LBD systems and their text mining and network analysis ingredients are explained in Section 3.7.