Analysis of Information Sources Based on the Cosine Similarity Algorithm, Affinity Propagation and Spectral Clustering Methods
摘要
This article explores the development of an optimized framework for analyzing and comparing information sources by leveraging large volumes of textual data through natural language processing (NLP) techniques. The study focuses on Telegram news channels as the primary sources of textual information. Text pre-processing steps, including cleaning, tokenization, and lemmatization, were performed to create a global dictionary comprising unique words across all sources. For each channel, a vector representation was constructed, where the dimensionality corresponds to the number of unique words in the dictionary, and the frequency of each word in the channel’s texts was represented in the respective vector positions. By applying the cosine similarity algorithm to these vectors, a square similarity matrix was generated, illustrating the degree of resemblance between different channels. Temporal analyses were conducted by evaluating the similarity of channels within specific time intervals, uncovering patterns and shifts in their informational strategies. Model parameters were fine-tuned to maximize differentiation among channels, thereby enhancing the analytical precision. Furthermore, clustering algorithms were employed to group channels based on their lexical similarity, providing insights into the formation of distinct information clusters. The findings highlight the effectiveness of the proposed approach for quantitatively assessing text similarity and clustering data from diverse sources. This method offers a robust tool for analyzing information channels, identifying interrelations between sources, tracking the evolution of their content strategies, and evaluating the socio-cultural implications of media outputs.