A scalable framework for scholarly similarity and novelty measurement with LLM-derived semantic embeddings
摘要
We develop a scalable LLM-embedding-based framework for semantic analysis of large-scale scholarly corpora. In this framework, modern LLM-derived and pretrained scientific-domain embedding models transform paper abstracts into dense semantic representations, while approximate nearest-neighbor (ANN) retrieval enables efficient corpus-scale similarity computation. Using 307,215 statistics-related papers, we compare six embedding models, including text-embedding-3-small and text-embedding-3-large from OpenAI together with scientific-domain and contrastive models. We evaluate whether Annoy, used as the scalability layer, preserves reliable semantic neighborhoods in these embedding spaces. We utilize the resulting embedding framework to construct a journal-level similarity network that reflects content-based proximity and introduce an LLM-embedding-based semantic novelty measure, defined by semantic distance from prior work. Computational experiments demonstrate semantic discrimination, retrieval fidelity, scalability, and robustness of the proposed framework. This approach supports large-scale scientometric analysis and provides a reproducible semantic perspective on scientific knowledge organization.